How many users do you need for usability testing?

Five to eight people per distinct user group finds most of the problems that stop somebody completing a task. Past that you are paying to watch the same issues again. If you have two genuinely different kinds of user, run five to eight of each rather than ten mixed.

The number people quote is five, it comes from a paper published in 1993, and it is right often enough to be useful and wrong often enough to waste a research budget. What follows is when five is genuinely enough, and the three situations where it is not.

How many users do you need for usability testing?

Five to eight people per distinct user group finds most of the problems that block somebody completing a task. Past that you are mostly paying to watch the same issues happen again. If you have two genuinely different kinds of user, run five to eight of each rather than ten mixed.

The reasoning behind the number is not magic. Blocking usability problems are common by definition: if a button is in the wrong place, most people looking for it will struggle. Common problems show up in the first few sessions, and each additional participant finds fewer new ones than the last. The curve flattens, and the flattening is the whole argument.

How much a session adds, participant by participant
% OF BLOCKING ISSUES SEEN SO FAR 0% 25% 50% 75% 100% 1 2 3 4 5 6 7 8 9 10

Illustrative, not a measurement of your product. The shape is what people mean by 'five is enough': the first participant finds a third of it and the tenth adds a point. It only holds inside one user group, which is the part usually left out.

Growth is a process. Nothing here happens overnight, and anybody promising you overnight results is lying to you.

Run a session from $100

When do you need more than five?

Three situations. When you have distinct user groups who use the product differently. When you are measuring something rather than finding problems. And when the decision is expensive enough that being wrong costs more than the extra sessions.

The first is the one that catches people. A tool used by both an administrator and an end user is two products wearing one interface, and five sessions split across both groups is really two and a half sessions of each, which finds almost nothing reliably. Count your groups first, then multiply.

The second is a different activity with the same name. Watching five people struggle tells you what is broken. It does not tell you that 43% of users struggle, and any number produced from five people is a number you should not put in a slide. If you need a percentage, you need a survey or an analytics comparison, not more sessions.

What is the difference between moderated and unmoderated testing?

In a moderated session a researcher is present and can ask why. In an unmoderated one the participant records themselves against a task list. Moderated finds out the reason behind the behaviour; unmoderated gets you more sessions for the same money and no follow-up questions.

The practical rule: use unmoderated when you know what you are looking for and want it confirmed across more people. Use moderated when you do not know what is wrong. The most valuable thing in a research session is usually the answer to a question the researcher decided to ask twenty seconds earlier, and that opportunity does not exist in a recording.

Choosing between the two
SituationModeratedUnmoderated
Nobody knows why people drop out Yes. The why is the deliverable. You will see where, not why.
Checking a redesign against the old one Useful, and expensive for the volume you want. Yes. More sessions, same budget.
Complex or unfamiliar product Yes. Participants get stuck in ways only a person can unstick. Sessions end early and you learn nothing.
Testing on a specific device or setup Possible, awkward to arrange. Yes, if recruitment screens for it.

Who should you recruit for user testing?

People who match the person you are trying to serve, screened on behaviour rather than on demographics. "Has bought something online in the last month" is a useful screen. "Aged 25 to 34" almost never is, because age does not predict how somebody uses a checkout.

The temptation is to recruit from your own audience, because it is free and fast. It also selects for the people who already worked out how to use your product, which is the group least able to show you why other people cannot. Existing users are the right group for questions about depth and retention. They are the wrong group for questions about first use.

  1. Screen on behaviour, not demographics What they do predicts how they will use the product. What year they were born does not.
  2. Give them the job, not the steps "Get a refund for the wrong-size item" produces real behaviour. "Click Returns, then choose a reason" produces a person following instructions.
  3. Say nothing while they struggle The silence is unbearable and it is the data. Every hint you give is a finding you have just deleted.
  4. Watch what they do before what they say People are unreliable narrators of their own behaviour, and reliably polite about products in front of the person who made them.
  5. Write the finding, not the anecdote "Four of six looked for export under Settings first" is actionable. "Users found it confusing" is a feeling with a number bolted on.

Is user testing worth it for a small product?

Usually more worth it, not less, because a small product has fewer places for a problem to hide and less traffic to reveal it statistically. If you have a thousand visitors a month, you will never A/B test your way to an answer. Five sessions will find it in an afternoon.

This is the part people get backwards. Testing is often treated as a thing you graduate to once you have a research budget. In practice, low traffic makes qualitative work more valuable, not less, because the quantitative alternative needs volume you do not have. The smaller you are, the more you should be watching individual people.

Want this done rather than explained?

Every service on the storefront is priced in the open in US dollars, with entry tiers at $100 so you can see the output before you commit to a programme. See what this costs, or read the questions people ask before they buy.