Theresa Wilkinson

Articles & Publications

Can AI Design a Good UX Study?

I started this project with a fairly simple question: Could AI design a good UX study? Not just generate a list of research questions or spit out some usability tasks. I wanted to know whether AI could approach research planning the way an experienced UX researcher would—look at the situation, make judgments, recognize what mattered, and create something that could actually be used on a real project.

To find out, I gave ChatGPT, Claude, Gemini, and Copilot the exact same assignment: design a qualitative usability study for a veterinary appointment scheduling application.

Each received the same scenario and prompt. I asked them to develop research questions, participant criteria and sample-size recommendations, and realistic usability tasks. I evaluated what they produced for completeness, accuracy, quality, practicality, and the amount of editing required.

My First Surprise: They Could All Do It

At least technically. All four produced usable starting points. They understood the assignment, identified the major pieces of a UX study, and gave me something I could work with.

But none of them produced a complete research plan I would have used without reviewing and changing it.

That distinction became important. I wasn't really interested in whether AI could produce something that looked like a research plan. I wanted to know whether it could make the kinds of decisions that go into a good one.

That's where things started getting interesting.

Research Questions: Strong, But Not Complete

At least technically. All four produced usable starting points. They understood the assignment, identified the major pieces of a UX study, and gave me something I could work with.

But none of them produced a complete research plan I would have used without reviewing and changing it.

That distinction became important. I wasn't really interested in whether AI could produce something that looked like a research plan. I wanted to know whether it could make the kinds of decisions that go into a good one.

That's where things started getting interesting.

Research Questions: Strong, But Not Complete

The research questions were actually one of the stronger parts of the AI-generated plans. The models identified many of the things I would expect a researcher to consider, although the depth and coverage varied considerably.

ChatGPT generated the most comprehensive research questions, covering the primary workflows as well as accessibility, digital literacy, terminology, navigation, error recovery, reminders, and trust. But there was a gap I noticed immediately.

None of the four models identified readability or comprehension as something that should be evaluated.

The scenario specifically said that the application served adults with varying levels of digital literacy. The models recognized digital literacy as an issue, but they didn't make the next connection I would make as a researcher: Can these users understand what the application is telling them?

That was one of the first places where I found myself doing more than editing AI output. I was applying my own research experience to identify something that wasn't there at all.

Then the Assumptions Started Appearing

Participant recruitment produced another surprise. Claude gave me the most comprehensive participant criteria, but recommendations varied considerably across the four models. Only one recommended a broad adult age range, and none fully addressed diversity, accessibility, and platform experience.

I also started noticing assumptions being introduced into the research plans.

For example, some models required participants to have previous digital scheduling experience. That wasn't necessarily a bad recruiting criterion—but it wasn't something I had specified.

That bothered me because it changes the study population. What about someone who owns a pet but doesn't already schedule appointments digitally?

The AI wasn't simply filling in my research plan. It was making decisions about my participants. That meant I had to evaluate not only what the AI recommended, but also what it had assumed about the users and why.

Where Was the Analysis?

One of my biggest frustrations was something I hadn't anticipated: the AIs didn't really analyze the information they generated.

I expected them to step back and look across the research questions, participant criteria, and usability tasks—to group similar items, count how often themes appeared, identify patterns, and quantify areas of agreement and disagreement. They didn't.

Instead, the responses tended to generalize: a model covered the primary workflows, included accessibility, or provided comprehensive participant criteria. But I wanted to know how many models identified the same thing, what they consistently missed, and where their recommendations actually differed.

As a researcher, that's what I would do with this information. I wouldn't just describe it. I would organize it, count it, compare it, and look for patterns. That analysis still had to be done by me.

The “Grad School” Problem

Then I got to sample size. ChatGPT recommended 8–12 participants and one round of testing. Copilot recommended 20–48 participants and four to six rounds.

Across the four models, the recommendations ranged from 8 to 48 participants and one to six rounds of testing.

This captured something that had been frustrating me throughout the exercise.

I had hoped the AI would approach the research more like an experienced UX researcher—looking at the situation, making judgments, and thinking about what would actually work on a real project.

Instead, some of the responses felt more like a grad-school exercise: thorough on paper, but not always practical. Four to six rounds of iterative testing may sound wonderful. Now give me the budget, participants, staff, and schedule to actually do it.

Real UX research involves tradeoffs. The theoretically ideal research plan isn't necessarily the right research plan for the organization, product, budget, or timeline. That's a judgment call. And I wasn't seeing enough of that judgment.

Accessibility Was Another Interesting Test

All four models generated usability tasks, but their coverage varied considerably. ChatGPT created the most comprehensive set. Copilot produced detailed, realistic multi-step scenarios but omitted accessibility testing and several secondary tasks.

What made that particularly interesting was that WCAG 2.2 compliance was explicitly stated in the scenario. Yet accessibility-related tasks still appeared inconsistently across the models.

That reinforced something I'd already begun to see: an AI can recognize a requirement without necessarily carrying that requirement all the way through the research plan. A researcher still has to notice that disconnect.

So, Can AI Design a Good UX Study?

Yes—with a fairly large asterisk. The AIs were useful. They generated ideas quickly, identified core research components, and gave me solid starting points. But they didn't replace the thinking involved in research planning.

I still had to ask:

And there was an unexpected tradeoff: AI reduced the time I spent drafting, but increased the time I spent reviewing and refining. I went into the study wondering whether AI could design a UX study. I came out of it with a slightly different question:

Can AI think about a UX study the way an experienced researcher does?

After Study 1, my answer was: not yet. But that made me curious about what would happen if I gave AI something harder to do.

View LinkedIn Version

Back to Portfolio

Contact

Email: theresaw@columbus.rr.com

LinkedIn: theresa-wilkinson