Research methods
Usability testing: how to plan, script and run a test, including with B2B users
How to run a usability test: write realistic tasks, script the session, choose moderated or unmoderated, test five users a round, and adapt for expert B2B users.
A usability test is a session where you give someone realistic tasks to complete with your product or prototype and watch what they do. Nielsen Norman Group describes three core parts: a facilitator who gives the tasks and asks follow-up questions, the tasks themselves, and a participant who is a realistic user (Usability Testing 101). Participants are usually asked to think out loud as they work.
Usability testing tells you whether people can do the job with what you built. It does not tell you whether they want it. That question belongs to customer interviews and the rest of product discovery.
A worked example
Suppose you are building accounts-payable software for restaurant groups, and you have redesigned the screen where a finance manager approves supplier invoices before payment. You want to know whether finance managers can spot an invoice that doesn't match its delivery, hold it, and approve the rest. This example is hypothetical, and we will use it throughout.
Decide what you are testing and with whom
Write the research question in one line: "Can finance managers at multi-location restaurant groups approve a weekly invoice batch and hold the exceptions without help?" Then decide who counts as a realistic user. NN/g says participants should be realistic users of the product: current users, or people with a similar background and the same needs.
For a typical qualitative study of one user group, NN/g recommends five participants. Jakob Nielsen's argument is that five users find most of the common problems, and that a budget for 15 users is better spent on three rounds of five, fixing problems between rounds (Why You Only Need to Test with 5 Users). If you have distinct groups of users, he suggests three to four from each of two groups, or three from each when there are three or more. That guidance is for finding problems. If you need metrics such as task success rates for a benchmark, you need more people; NN/g's follow-up on quantitative studies recommends 20.
In the example, the controllers who approve payments and the general managers who send invoices in are different groups. If both use the screen, test three or four of each.
Write tasks that set a goal without giving the answer
Tasks are where many tests go wrong. GOV.UK's guidance says good tasks set a clear goal, are relevant and believable to the participant, are hard enough to reveal problems, and don't hint at how to complete them (Using moderated usability testing). NN/g warns that small wording errors can confuse participants or nudge them toward an answer.
| Weak task | Better task |
|---|---|
| Click "Hold" on the invoice with a price mismatch | This week's produce invoice from one supplier looks higher than usual. Decide what to do with it, then pay everything else that is ready. |
| Use the filter to find invoices over $1,000 | Your owner wants to see the biggest bills before they go out. Find them. |
The weak versions name the button. The better versions describe the situation a finance manager would actually be in.
GOV.UK also recommends real data when you can handle it securely: participants using their own documents are more engaged and reveal more than those using dummy data. If you must use dummy data, make it look like theirs, with supplier names and amounts a restaurant group would recognize.
Script the session
GOV.UK suggests sessions of 30 to 60 minutes and a discussion guide that holds your introduction, the tasks and a planning checklist. A simple outline:
- Introduction (about 3 minutes). Explain what will happen, confirm consent and recording, and say you are testing the product, not them.
- Warm-up (about 5 minutes). Ask about their job and how they approve invoices today. Use their answers to make the tasks feel like their week.
- Tasks (30 to 40 minutes). Read each task aloud, or have the participant read it aloud. Ask them to think out loud, then stay quiet. Step in only if they are completely stuck or say something you need to understand, using neutral questions such as "What makes you say that?"
- Wrap-up (5 to 10 minutes). Ask about anything you saw but didn't understand, then any final thoughts.
How to ask for recording consent covers the consent step, and how to take interview notes covers capturing what happened.
Choose moderated or unmoderated
In a moderated test, a facilitator is present, in person or on a video call, and can ask follow-up questions. In an unmoderated test, the participant completes the tasks alone using a testing tool that records the session for you to watch later (Remote Usability Tests: Moderated and Unmoderated).
| Moderated | Unmoderated | |
|---|---|---|
| Follow-up questions | Yes, specific to what the person did | Only the same pre-written questions for everyone |
| If the participant gets stuck | You can help them recover | No real-time help; you find out afterward |
| Thinking aloud | You can prompt quiet participants | No one is there to remind them |
| Best for | An overall review of a whole workflow | A few specific elements or a small change; tight timelines |
| Written instructions | Can be clarified live | Must stand on their own, so pilot them first |
NN/g also advises recruiting a few extra people for remote studies, because no-show rates can be higher and you won't know whether an unmoderated session is usable until you watch it.
For a B2B workflow like invoice approval, moderated sessions are usually the better choice. The tasks depend on context, and you will want to ask why someone hesitated.
Adjust for expert and B2B users
Professionals who use a tool every day are harder to test. In Testing Expert Users, Nielsen says the basics stay the same (representative users, realistic tasks, thinking aloud) but several things change:
- Give harder, deeper tasks. Experienced users can work through larger, more realistic problems.
- Expect less narration. Skilled behavior is often automatic, so experts may not be able to say why they did something. Watch what they do more than what they say.
- Steer critics back to the task. Experienced users may turn into design reviewers. Accept their comments, then bring them back to using the product; behavior is the data you came for.
- Recruit typical users, not stars. If account managers help you recruit, ask for average users rather than the top performers managers like to show off.
- Build experience for a new product. If nobody has used it yet, Nielsen suggests training participants quickly or giving them practice time before the test.
B2B participants are also busy and hard to schedule. Keep sessions short, book around their calendar, and pay for their time. Research interview budgets covers what to pay, and how to interview B2B users, buyers and champions covers who does what inside a customer.
Analyze, fix and test again
After each session, note what happened on each task: completed, completed with difficulty, or not completed, and where it went wrong. Look for problems more than one participant hit. Fix the worst ones and test again with new people. Nielsen's case for small rounds depends on that second round, which checks whether your fixes worked and often finds deeper problems the first round's surface issues hid.
Your next step
Write one research question, three tasks that describe situations instead of buttons, and a one-page guide. Book five sessions with people who do the job.
If your users are professionals you can't reach through your own customer list, Instant Expert can find people who match a description you write, such as "finance managers at multi-location restaurant groups." You review who it finds, it sends your invitations, and you pay only for calls that get booked. The finance professionals in restaurants directory page is one place to start.