# Can AI replace user research? It cannot meet Awad.

My own design engine recommended a near clone of the claude.ai landing. What AI does well in research, where it fails, and why UX careers will thrive.

https://laithjunaidy.com/writing/can-ai-replace-user-research
By Laith Aljunaidy. Published: 2026-10-05.

---

![Four landing pages generated by ux-skill from four briefs: a family clinic, a developer tool, a family restaurant, and an Arabic invoicing app](https://laithjunaidy.com/writing/design-process/uxskill-4-brands.jpg)

*Four briefs through [ux-skill](https://laithjunaidy.com/research/ux-skill), four different systems. On its own homepage brief, the same engine first handed back the most common answer in the category.*

In May 2026, [ux-skill](https://laithjunaidy.com/research/ux-skill) needed a new homepage. ux-skill is the design engine I build. Its whole reason to exist is to catch the fingerprints that make AI-built interfaces look the same. So we did the obvious thing and ran our own recommender on our own brief: a developer tool landing, tone editorial, confident, calm, cinematic. Inter as display, purple to blue gradients and three equal cards were on the forbidden list.

Twelve seconds later the engine answered. Cormorant Garamond on cream, with a warm sienna accent. Which is, almost exactly, the claude.ai landing.

The engine had obeyed every rule in the brief. It read "editorial" and "calm" and reached for the highest-similarity match, which was the look the AI tools in the category had already shipped. We filed [two bugs against ourselves](https://uxskill.laithjunaidy.com/blog/dogfooding-design-engine.html). The editorial tag over-weighted one family of book serifs. The calm tag had been coded as "light canvas required", so near-black palettes never got a chance. After the fixes the same brief returned a charcoal canvas, one amber accent and Fraunces, a system none of us would have picked on taste alone.

ux-skill makes no language-model calls. It is a deterministic scoring engine, and it still did what a model does with a familiar question. It gave the most likely answer. The most likely answer is the average of what is already out there.

Keep that picture in mind, because it is the whole argument about AI and user research. Ask a model to play your user and you get the most likely user. Research exists to find the person who is not.

## Key takeaways

- AI cannot replace user research. It can draft the script, transcribe and translate the sessions, sort the notes, list the failure states, and write the first pass of the report.
- A synthetic user is the average of what a model has read. Nielsen Norman Group found synthetic users are sycophantic, care about everything equally, and report experiences they never had.
- The narrower your user, the worse the simulation. NN/g's own example of where not to use synthetic users is a rice farmer in Vietnam. A livestock breeder in the Gulf is the same case.
- The UX job market is stabilising, entry level is still hard, and the value is moving from producing screens to judgement, research, systems and taste.
- Use AI the way NN/g describes it: like an intern. Give it context and constraints, then check every line against a real quote.

## What did ten breeders tell us that a model could not?

In 2023 I was designing [Wfrah](https://laithjunaidy.com/work/wfrah), an app for sheep and goat breeders in the Gulf. Before any screen existed, the founder, Abdullatif, told us what we were up against. Breeders run the trade the way their fathers taught them, and they do not believe an app knows their animals better than they do.

So we went and listened. Ten in-depth interviews with breeders of different experience, plus time watching the tools they already used. Ask a breeder why the farm loses money and he gives you the same answer every time: feed is expensive, and the buyers pay too little.

That is the answer people say out loud, and the answers people say out loud are the ones that end up written down. Anything trained on what people write would hand it straight back to you.

The interviews showed what sat underneath. Many breeders could not work out their herd's daily nutritional needs from the equations, and the belief that the market was to blame made them overlook nutrition and management. The most useful thing in the app became a [nutrition calculator](https://laithjunaidy.com/writing/designing-wfrah) that does the maths for each animal and each growth stage. Nobody asked for it. The interviews showed how many needed it.

![Awad, the fictional breeder persona from the Wfrah research, crouching beside a ram in a pen](https://laithjunaidy.com/writing/wfrah/persona.jpg)

*Awad. A fictional persona built from ten real interviews: proud of his family's methods, curious whether there is a better way.*

The research became one persona, Awad. He is proud of doing things the way his family has for generations, frustrated with how hard the herd is to manage, and curious whether there is a better way. That tension, pride and curiosity at once, became the whole design brief: an app that lectures Awad loses him, and an app that shows him something useful on the first screen might keep him.

Awad is fictional, built from ten real interviews. That is the difference between a persona and a synthetic user, and the rest of this piece hangs on it.

![Hisham, the mobile developer, and me working on Wfrah in the same room](https://laithjunaidy.com/writing/wfrah/team.jpg)

*Working with Hisham, the mobile developer, on the Wfrah build.*

## What does AI do well in user research?

More than its critics admit. My [design process guide](https://laithjunaidy.com/writing/design-process-checklist) has a "with Claude" paragraph under many of its methods. The pattern in all of them is the same: the model drafts, sorts and lists. A person decides.

Nielsen Norman Group's guide to [accelerating research with AI](https://www.nngroup.com/articles/research-with-ai/), last reviewed in January 2026, lands in the same place. It finds AI most helpful in the planning and analysis stages, and weakest in the middle, where someone has to watch a person use a product.

| Research task | What AI does well | What you still check |
| :- | :- | :- |
| Desk research | Starts the reading and gathers sources | Made-up facts and sources. Open every one |
| Interview scripts and test tasks | Drafts many options fast | Leading questions and "would you" questions |
| Consent forms, screeners, plans | Fills your template for this study | Wrong permissions, wrong details |
| Transcription and translation | Speakers, timestamps, summaries | Errors on accents, and in some languages |
| Clustering notes | A first pass at themes | Groups built on keywords, an overfull "other" pile |
| Failure states | Lists every way a step can break | That each state names the field and the fix |
| Report drafts | Grammar, tone, plain words for stakeholders | Details that never came from the research |

**Drafting the script.** A blank page is the slowest part of planning an interview. I paste the research question and ask for a 30 minute script about the last time this person did the task, with no hypothetical questions, no questions that name a feature, and a flag on any question that suggests its own answer. Then I delete every "would you" question it slipped in anyway. NN/g saw the same thing with test tasks: AI often writes leading or priming tasks, even with good examples and explicit rules. The draft saves an hour. The edit is still yours.

![User interview script and notes template: 18 questions in five sections beside a notes sheet with quotes, observations and follow ups](https://laithjunaidy.com/store/user-interview/en-light.jpg)

*The free [user interview template](https://laithjunaidy.com/store/user-interview), filled with a worked example: Bunn, a fictional coffee app. Script on the left, one notes sheet per participant on the right.*

**Transcribing and sorting.** After the sessions, the model earns its keep. I paste the transcripts and ask for every observation as one line with the participant ID and an exact quote, then for proposed groups, with one rule: nothing said by only one participant becomes a theme. Sorting is what models do well. Filling the gaps with plausible findings is what you have to catch. NN/g puts it plainly: AI clusters often group by keyword, and many notes end up in "other".

![Affinity map template: six interviews sorted into five themes for the Bunn coffee app](https://laithjunaidy.com/store/affinity-map/en-light.jpg)

*An [affinity map](https://laithjunaidy.com/store/affinity-map) from the Bunn example. Six interviews, five themes, and every note carries its participant tag. The tag is what lets you check the model.*

**Listing failure states.** This is where AI is better than most designers on a deadline. Give it a happy path and ask for every failure at every step: invalid input, timeout, offline, a double tap, permission denied, empty data. Then check that each state tells the person what went wrong and how to fix it. On the [Dot](https://laithjunaidy.com/research/dot) counter, the line for a short phone number reads "Nothing added. Phone is short by 1." A model will list that state. The rule that every failure starts with "Nothing added", because a cashier's first question after an error is whether it went through, came from designing for the person at the counter.

## Where does AI fail at user research?

At the one job research exists for: finding out what is true about people you have not met. The failure has a name. A synthetic user is an AI-generated profile that imitates a user group and answers interview questions as if it were one of them. Products sell this as research without the users.

Nielsen Norman Group [tested synthetic users](https://www.nngroup.com/articles/synthetic-users/) against three studies they had already run with real people. Here is what they found.

**It wants to please.** Real learners in an online training study admitted they started courses and did not finish them. The synthetic users said they had completed every course they mentioned. Real participants mostly ignored discussion forums because the interactions felt contrived. The synthetic users praised the forums at length. NN/g's explanation is that the model is trained on literature about why forums are good for learning, and good for you is not the same as used by you.

**It cares about everything.** Asked what makes a course engaging, a synthetic user listed seven factors with equal enthusiasm. Real people care about some things far more than others, and that ranking is the point of research. A list of seven equal needs cannot tell you what to build first.

**It remembers things that never happened.** A model has not used your product, so its stories about past experience are invented. NN/g notes the profiles are more accurate in their opinions than in their stories, which is backwards for research, where the story is the evidence. When they pitched a drone delivery idea to a synthetic medical sales rep, it called the idea a possible game-changer.

**It is flat.** NN/g's sharpest line is that synthetic users feel like a flat approximation of the experiences of tens of thousands of people, because they are. That is my homepage story again. The engine did not fail at maths. It returned the middle of the cluster.

The narrower your user, the worse it gets. NN/g's own advice is to avoid synthetic users for niche or specialised populations, and its example is a rice farmer in Vietnam, because the model's data on that person is patchy. Awad is that case exactly. A breeder in the Gulf who learned the trade from his father did not write much of the text a model learned from. Neither did a cashier at a busy counter.

**What people say and what they do.** The second gap is behaviour. Models are good with words, and research is often about the space between words and actions. In the Bunn example I use across my templates, Rana, the persona, calls the late delivery fee a trick in her interview. In the notes she checks the fee twice before paying, once on the menu and once at checkout. The quote and the behaviour together are the finding. A transcript only has the quote.

![Empathy map for Rana ordering coffee for a 10:00 meeting: says, thinks, does and feels, with pains and gains](https://laithjunaidy.com/store/empathy-map/en-light.jpg)

*The [empathy map](https://laithjunaidy.com/store/empathy-map) from the Bunn example. Says and does sit side by side, because the finding usually lives where they disagree.*

NN/g says the same about usability tests: no human or AI tool can analyse a usability session from the transcript alone, because people do not say every click, and often do one thing while saying another. Its 2026 guidance still advises against letting AI moderate usability tests. Rana's double check of the fee is a behaviour. You only get it by watching.

**It cannot weigh the source.** A good researcher asks questions a model does not: did the interviewer prime this answer, was the participant embarrassed, did this person even fit the recruit. NN/g lists exactly these as context-informed judgement beyond current tools.

## What does the evidence say?

The people who study this agree on the shape: AI is everywhere in research work, and almost nobody trusts it unchecked.

| Source | Finding |
| :- | :- |
| [User Interviews, State of User Research 2025](https://www.userinterviews.com/state-of-user-research-report), 485 researchers | 80% use AI, up 24 points in a year |
| | 91% worry about accuracy and hallucinations |
| | 63% fear AI could devalue human insight and critical thinking |
| [Figma, 2025 AI report](https://www.figma.com/blog/figma-2025-ai-report-perspectives/), 2,500 designers and developers | 78% say AI significantly enhances their efficiency |
| | 32% say they can rely on AI output |
| | 69% of designers are satisfied with AI tools, against 82% of developers |
| | 31% of designers use AI in core design work, against 59% of developers in core development |
| [MeasuringU, ChatGPT in tree testing](https://measuringu.com/chatgpt4-tree-test/), 33 people, 10 tasks | People succeeded on about 51% of tasks. ChatGPT found the right path on all ten in one run and nine of ten in four others |

Read the Figma pair together. Seventy-eight percent feel faster, 32% trust the output. That gap is the job. Someone has to check the work, and the person who checks it has to know what true looks like.

The MeasuringU study is the cleanest proof of the averaging problem. Asked to find items in the IRS website's navigation, ChatGPT found them far more often than people did. A participant who never gets lost is useless in a test built to find where people get lost. Its ratings of how hard each task would feel came reasonably close to what people reported. It could predict the opinion. It could not reproduce the behaviour.

NN/g's verdict: use synthetic users, if at all, for desk research and hypotheses before a real study. Never present their output as findings, and be most careful where research is already hard to get, because that is where a cheap imitation replaces the real thing.

## Is AI taking UX jobs?

Not the way the 2023 headlines promised. Nielsen Norman Group's [State of UX 2026](https://www.nngroup.com/articles/state-of-ux-2026/) describes the arc: a hiring boom after the pandemic, a sharp drop from 2023 into 2024 as budgets tightened, and stabilisation from late 2024 through 2025. It also names the story that made the drop worse: AI hype suggested new tools could rapidly replace designers and researchers. NN/g says that was not true, but it was convenient for cost cutting.

The recovery is uneven. Senior practitioners and generalists are coming back faster. Entry-level roles remain scarce and highly competitive, and NN/g expects 2026 to stay that way. The User Interviews survey shows the same pressure from the inside: 49% of researchers feel negative about the future of the field, up 26 points, and 67% are negative about career opportunities.

It also shows what the headlines miss. Twenty-one percent said their company laid off researchers, about the same as in 2024, and researchers were not hit harder than design or engineering. Thirty-five percent were promoted or took on more responsibility, and 20% moved into a new function. The report's own reading is that research may be turning from a role into a skill that other roles need. Fewer people may hold the title. More people need the work.

So the honest answer is that AI did not delete the job. It raised the bar for the job, and the bar is highest at the door.

## How is the UX job changing?

The work moves from producing screens to deciding what the screens should be. NN/g puts it bluntly: UI is no longer a differentiator. Design systems made interfaces cheaper to produce, and AI tools amplify that. In their words, if you are just slapping together components from a design system, you are already replaceable by AI. What is not easy to automate: curated taste, research-informed context, critical thinking and careful judgement.

I watch this happen in my own work. On the Dot counter, the screens were never the hard part. The first version asked for the amount first, then a country, then the phone, then a button: six taps plus typing, with nothing in focus when the screen opened. Any tool can draw that screen. The rebuild started from one line in the spec: a tired cashier at a busy counter, one hand on the phone and one on the cup. Phone first, the amount field takes focus by itself, and a returning customer now costs two taps.

The value moved upstream, into a sentence about a person. My [design process guide](https://laithjunaidy.com/writing/design-process-checklist) says it this way: AI makes building cheap, so the expensive mistake moves upstream, to building the wrong thing quickly. Four shifts follow from that.

**From screens to systems.** When a design system and a model can assemble a screen, the designer's work is the system: tokens, patterns, the rules for every state. That is why ux-skill compiles a fresh system from each brief, and why I wrote that [brand specs are training data, not templates](https://laithjunaidy.com/writing/brand-specs-are-training-data).

**From output to judgement.** NN/g's piece on [design taste in the era of AI](https://www.nngroup.com/articles/taste-vs-technical-skills-ai/) argues that technical capability does not equal taste, and that selection and discernment are what make work stand out once anyone can produce something. My homepage bug is the proof in miniature. The engine produced a perfectly competent system. Knowing it was wrong for us took someone who could see the category it had fallen into.

**From the average user to the specific one.** The defaults are converging. When we [read 150 shadcn/ui repos](https://laithjunaidy.com/writing/taken-apart-shadcn-defaults), the most-starred projects ran on one of two typefaces and a grey primary. The designer worth hiring is the one who can say who this is for, in detail, with evidence.

**From English first to local context.** This is where the gap is widest, and where I work. When we [checked 179 funded MENA startups](https://laithjunaidy.com/writing/taken-apart-arabic-100), 44% had no Arabic website. A persona translated from an English template records what the English author expected, and a model that learned mostly from English text has the same problem. The Arabic version of my free [persona template](https://laithjunaidy.com/store/persona) was written in Arabic for that reason, with a bio that reads like someone from Amman. NN/g's niche-user warning applies at country scale: the less a population writes online, the less a model knows it.

## Why will the UX career thrive?

Because every force in this story raises the price of the thing only people can do. Screens get cheaper, so knowing which screen matters more. Building gets faster, so building the wrong thing gets more expensive. AI features multiply, and NN/g expects trust to be a major design problem in 2026, which is a research problem first. In Figma's survey, 52% of AI builders said design matters more for AI products than for traditional ones, and 95% said at least as much.

What that means for a career, concretely:

**What to learn.**

- **Interviewing and observation.** The two skills a model cannot perform. Run the [user interview](https://laithjunaidy.com/store/user-interview) yourself, ask about the last time, and watch people do the task where they normally do it.
- **Synthesis with traceability.** An [affinity map](https://laithjunaidy.com/store/affinity-map) where every note carries a participant ID. It is also how you check the model's sorting.
- **Usability testing.** Watch five people. NN/g's guidance is still to keep AI out of moderating these. The free [usability test template](https://laithjunaidy.com/store/usability-test) logs every finding with its evidence quote.
- **Systems.** Tokens, components, states, handoff.
- **Your local context.** The language, the counter, the pen, the phone in bad light. The narrower your users, the more this knowledge is worth.
- **Working with AI.** More than 80% of designers and developers in Figma's survey say learning to work with AI will be essential. Learn to brief it, constrain it and check it.

**What to stop doing.**

- Treating a pile of screens as the deliverable. Ship the decision and the evidence behind it.
- Asking a model to play your user, then designing for its answers.
- Running "would you use this" interviews. People and models are both happy to say yes.
- Presenting AI summaries as findings. NN/g calls that unethical when the source is synthetic, and it damages how the organisation sees research.

![Usability test template for Bunn: plan, five tasks with pass, partial and fail results, and findings sorted by severity with evidence quotes](https://laithjunaidy.com/store/usability-test/en-light.jpg)

*The [usability test template](https://laithjunaidy.com/store/usability-test) from the Bunn example. Every finding carries the quote that proves it. That column is the part a model cannot fill.*

**How to use AI as a junior researcher you check.** NN/g says AI works best like an intern: ample instructions, context, constraints and corrections. Treat it that way.

1. **Brief it like a new hire.** The research question, the participants, what you already know. Without context, NN/g found the findings came back vague and unranked.
2. **Give it a template.** It makes fewer mistakes filling a good form than inventing one.
3. **Demand sources.** Every observation with a participant ID and an exact quote. Every desk research claim with a link you then open.
4. **Ban the gaps.** Tell it to leave a field empty when the notes do not support it. Then delete any line with no ID.
5. **Keep the choices.** Drafting three problem statements is the model's job. Choosing the problem is yours, and so is setting the severity of a finding. The model cannot know which delay costs money.
6. **Never let it be the user.** If you use a synthetic user at all, use it to rehearse your script or list topics before the real interviews, and label everything it says as a hypothesis.

The full set of prompts, method by method, is in the [design process guide](https://laithjunaidy.com/writing/design-process-checklist), and the nine free templates it uses are in the [store](https://laithjunaidy.com/store).

None of this is new. [Nobody wants your product](https://laithjunaidy.com/writing/nobody-wants-your-product), they want the outcome, and the outcome lives with a person you have to go and meet. AI made everything around that meeting faster. It did not make the meeting optional.

## FAQ

### Can AI replace user research?

No. A model drafts scripts, sorts interview transcripts and lists failure states well. It cannot tell you what Awad does in his pen, or why Rana checks the fees twice. Ask it to play your user and you get the average of what it read, which is exactly the assumption research exists to test.

### Will AI replace UX designers?

Not the ones who do the judgement part. Nielsen Norman Group's State of UX 2026 says designers who only assemble components from a design system are already replaceable by AI, and that taste, research-informed context, critical thinking and judgement are hard to automate. The job shifts toward those.

### Is UX a good career in 2026?

Yes, with a harder entry. NN/g reports the UX job market stabilised from late 2024 through 2025, with senior and generalist roles recovering faster than entry-level ones. In User Interviews' 2025 survey, 35% of researchers were promoted or took on more responsibility. Juniors who can interview, synthesise and test, and who know a local market, stand out.

### Can ChatGPT do user research?

It can do parts of it: desk research, drafting scripts, transcribing, a first pass at clustering, and report drafts. It cannot observe behaviour, and when it plays a user it tends to agree, to rate everything as important, and to invent past experiences. In a MeasuringU tree test it found the right answers far more often than real people, which made it useless as a participant.

### What are synthetic users?

AI-generated profiles that imitate a user group and answer research questions as if they were one of its members. NN/g recommends using them, if at all, only for desk research and hypotheses before real research, never as findings, and avoiding them for niche or specialised users.

### Can AI analyse usability tests?

Not reliably. NN/g says no human or AI tool can analyse a usability session from the transcript alone, because people do not say everything they do, and its 2026 guidance still advises against AI moderating usability tests. AI can transcribe the sessions and draft the findings list. You set the severity and check each quote.
