Theresa Wilkinson

Articles & Publications

I Asked Four AIs to Analyze the Same UX Survey. Only One Finished.

What happened after ChatGPT, Claude, Gemini, and Copilot received the same simplified synthetic dataset and the same instructions

I expected the difficult part of this experiment to be checking four sets of calculations. Instead, most of the models failed before there was much analysis to verify:

That was especially disappointing because analysis is my favorite part of research. I had been looking forward to getting four sets of results, digging into the data, comparing what each AI noticed, and seeing whether they found the same patterns I did. Instead, three of the four never made it far enough for me to do that.

What I tested

This was not a survey of 60 real UX professionals. The spreadsheet contained 60 synthetic responses created specifically to test whether AI tools could analyze UX survey data accurately and follow detailed instructions. My real survey currently has five respondents and was kept completely separate from this experiment.

The synthetic dataset contained 21 columns, including categorical questions, a multiple-selection question about AI tools, nine 1–5 trust ratings, and trust-most/trust-least questions.

Why I simplified the spreadsheet

The first version of the spreadsheet contained five tabs: one complete Data tab and four additional role-named tabs. All four AIs raised concerns about the multiple tabs, including ChatGPT. That made the spreadsheet structure a possible explanation for any problems.

To remove that variable, I deleted the extra tabs and created a simplified spreadsheet containing only the complete Data tab. I then resent the simplified file, the survey text, and the same analysis instructions. The underlying 60 synthetic responses did not change.

The results described below reflect what happened after I simplified and resent the spreadsheet. This matters because the later failures cannot be attributed simply to choosing among several worksheet tabs.

The instructions

Each AI received this direction:

Analyze the attached spreadsheet exactly as provided. Do not change, clean, standardize, recode, or reinterpret any responses. For each survey question, calculate the count and percentage for every response… Show all calculations clearly enough to verify.

The prompt also required descriptive statistics for rating questions, per-option counts for multiple-selection questions, clarification when a label was unclear, and no assumptions beyond the data.

What I was looking forward to

I love the analysis stage of research. This is where I get to play in the data: looking for patterns, questioning assumptions, checking whether an idea is actually supported, and finding the unexpected result that changes how I understand the problem.

I created the synthetic dataset so I would know what was in it, then planned to compare four independent analyses of the same responses. I expected to spend my time checking calculations, comparing interpretations, and deciding which findings were real. That was the part of the experiment I was most excited about.

What happened

ChatGPT: completed the analysis

After receiving the simplified spreadsheet, ChatGPT identified the Data tab as the source containing 60 responses and analyzed it without changing the submitted file. It calculated counts and percentages for each question, per-tool counts for the multiple-selection question, and descriptive statistics for all nine trust-rating questions.

It also flagged that the numeric trust ratings did not include verbal scale definitions in the Data tab instead of inventing labels or silently interpreting the values.

I manually checked the results. For example, the multiple-selection counts were ChatGPT 41, Gemini 30, Claude 18, and Microsoft Copilot 14—103 selections across 60 synthetic respondents. The counts matched.

Claude: Described more work than it delivered

Claude correctly recognized that the simplified spreadsheet contained one Data tab with 60 respondents and 21 columns. It then said that an Analysis tab contained counts, percentages, the per-tool breakdown, and statistics for all nine rating questions.

The ‘Analysis’ tab has every count/percentage as a live formula pointing at ‘Data’…

But the actual deliverable contained only one set of calculations—not the comprehensive analysis described in the response. Claude also created or returned a spreadsheet with an Analysis tab even though the prompt explicitly said not to change the spreadsheet.

Simplifying the spreadsheet helped Claude identify the file structure correctly, but it did not solve the larger problem: the written description of the work still did not match the work that was actually delivered.

Gemini: Continued analyzing the wrong dataset

I initially uploaded the wrong spreadsheet, then sent the corrected, simplified spreadsheet to all four AIs. That mistake unexpectedly created a useful stress test: Could each model recognize that a newer file had replaced the earlier one?

Gemini did not. It continued displaying results from the superseded dataset after the corrected spreadsheet was uploaded. I sent five prompts asking it to stop and explaining that it was using the wrong file, but it continued returning the old data.

I am unable to access or process the raw contents of AI UX Survey Data.xlsx directly…

After saying that spreadsheet-reading capabilities were unavailable, Gemini displayed data from the earlier spreadsheet anyway. It also did not respond to my request for a Word document. I abandoned the session because I could no longer trust it to identify which dataset was authoritative.

Removing the extra tabs did not correct Gemini’s use of the wrong file. The failure was file replacement and data provenance, not tab selection.

Copilot: Could not recognize the uploaded file

Copilot failed differently. Even after I resent the simplified spreadsheet, it repeatedly asked me to upload the spreadsheet. The page also stopped scrolling, making the exchange harder to review.

When I wrote, ‘Why are you freaking out?’ Copilot tried to explain the problem—but it treated my earlier complaint, ‘Why can’t I scroll up the page?’ as if it were an open-ended survey response.

The response you shared (‘why can’t I scroll up the page?’) appears to be a single open-ended comment…

Copilot had stopped distinguishing between my instructions, my interface complaint, and the research data. No analysis was completed. Simplifying the spreadsheet did not help because Copilot did not reliably recognize the file as an input.

Results at a glance

AI results table

Results at a glance: Only ChatGPT completed the requested analysis using the corrected spreadsheet.

What changed after removing the extra tabs

The simplified spreadsheet improved one part of the test: it removed ambiguity about which worksheet contained the data. Claude accurately described the new single-tab structure, and ChatGPT completed the analysis from that tab.

But, simplification did not make the other systems complete the task:

The extra tabs were therefore a usability issue worth correcting, but they were not the main cause of the observed failures.

By that point, I was not just frustrated with the tools. I was genuinely disappointed. I had expected four sets of analysis results to explore; instead, I spent the experiment troubleshooting uploads, correcting file confusion, and documenting failures. I never got the chance to play in the data the way I had planned.

The bigger finding: Analysis accuracy starts before the calculations

I designed this test to evaluate counts, percentages, multiple-selection breakdowns, and descriptive statistics. But three models exposed a more fundamental issue: before an AI can analyze research correctly, it must reliably identify the current file, distinguish data from conversation, follow preservation instructions, and accurately report what it completed.

A perfectly calculated percentage is still useless if it came from the wrong dataset. A polished summary is not evidence that the promised analysis exists. And a model cannot analyze a spreadsheet it does not recognize as an input.

For UX researchers, this means verification must begin with data provenance—not with the final numbers.

What researchers should verify

What this test does—and does not—show

This was an exploratory comparison of four individual sessions, not a controlled benchmark of every version or configuration of each product. The spreadsheet was revised during the process to remove four unnecessary tabs, and the results reported here are from the simplified resubmission. File handling and interface behavior may vary across sessions. The dataset was synthetic, and the results should not be interpreted as opinions held by 60 real UX professionals.

However, the observed failures were real within these sessions. They demonstrate why research analysis cannot be evaluated only by how polished the final response sounds.

Overall takeaway

Only one of the four AIs completed the requested analysis after receiving the corrected, simplified spreadsheet while preserving the submitted file. The others failed in different ways: incomplete work presented as comprehensive, continued use of an obsolete dataset, and failure to recognize the uploaded file at all.

Removing the extra tabs was an important methodological correction because it eliminated a plausible source of confusion. But it did not eliminate the major failures. The surprising lesson was that the largest risk was not a complicated statistical error. It was losing track of the research data before the analysis had even begun.

The experiment still produced a finding, but it was not the experiment I had hoped to conduct. I wanted the detective work: four analyses to compare, patterns to investigate, and unexpected findings to question. Mostly, I learned how easily the analysis can disappear before a researcher ever gets the chance to begin.

Open to contract UX research opportunities.

View LinkedIn Version

Back to Portfolio

Contact

Email: theresaw@columbus.rr.com

LinkedIn: theresa-wilkinson