# Usability Test Script with AI: A Step-by-Step Guide

> How to use AI to generate usability testing scripts: session structure, prompts for tasks and questions, results analysis. Ready-to-use templates and examples.
> Author: Roman Belov · Published: 2026-07-14 · Source: https://futurecraft.pro/blog/usability-test-script-ai/

85% of usability problems surface with just the first five test participants — a classic finding from [Nielsen Norman Group](https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/). The bottleneck isn't running the test — it's preparing the script. The average PM spends hours on it. AI cuts that to 30–40 minutes, and the quality is often better: the model covers scenarios systematically rather than from memory.

This article walks through how to generate a complete usability test script: from session structure to results analysis. With prompts you can use right now.

## Usability Test Script Structure

A usability test script has six blocks. Each one does a specific job. Skip any of them and your data quality takes a hit.

| Block | Duration | Purpose |
|-------|----------|---------|
| Introduction and consent | 3–5 min | Reduce anxiety, obtain recording consent |
| Pre-test survey | 3–5 min | Understand participant context and experience |
| Tasks (core) | 25–35 min | Observe behavior during task completion |
| Post-task questions | 5–7 min | Gather impressions after each task |
| General survey | 5–7 min | Assess overall product perception |
| Wrap-up | 2–3 min | Thank participant, ask what was missed |

Total session: 45–60 minutes. Go past an hour and participants tire — data reliability drops with them.

## Introduction: Prompt for Generating an Intro Script

The introduction sets the tone for the whole session. The participant needs to understand that the product is being tested, not them; that there are no right or wrong answers; and that thinking aloud helps, not hurts.

```
You are a UX researcher. Write an introductory script for a moderated
usability test.

Product: [name and brief description]
Format: [remote/in-person]
Recording: [yes/no]

The script must include:
1. Moderator introduction (name, role — no job title)
2. Explanation of the goal: we're testing the product, not the participant
3. Request to think aloud
4. Recording instructions and consent
5. Reminder: participation can be stopped at any time
6. Question "Any questions before we begin?"

Tone: calm, friendly, not corporate.
Length: 150–200 words.
```

The phrase "we're testing the product, not you" matters more than it sounds. Without it, participants try to please the moderator instead of working naturally. That skews your data.

## Pre-Test Survey: Participant Context

Before handing out tasks, you need to understand who you're working with. A participant's experience changes how you interpret their behavior. A novice stuck on navigation and an expert stuck on the same navigation signal completely different problems.

```
Generate 5–7 pre-test questions for a usability test.

Product: [name]
Product category: [e.g., task manager, travel app, CRM]
Target audience: [description]

Questions should reveal:
- Experience using similar products
- Frequency of solving the problem the product addresses
- Current approach to solving this problem (without our product)
- Level of technical proficiency (indirectly, no direct question)

Format: open-ended questions. No yes/no answers.
Don't mention competitors by name.
```

What you get from this block: a participant profile you'll use during analysis. "Participant 3 uses similar products daily but couldn't find feature X" tells you more than "Participant 3 couldn't find feature X."

## Generating Tasks: The Core of the Usability Test

Tasks determine what you're actually testing. A bad task produces useless data. A good task hits three criteria: it's realistic, specific, and hint-free.

Bad task: "Find the filter button and filter results by price." That's a hint. The participant knows to look for a button. In real life, they might not even know filters exist.

Good task: "You're looking for a hotel in Bangkok for 3 nights. Budget — no more than $80 per night. Show me how you'd find something."

```
Generate a set of 5–8 tasks for a usability test.

Product: [name and description]
Key flows to test:
1. [flow 1, e.g.: new user onboarding]
2. [flow 2, e.g.: creating the first project]
3. [flow 3, e.g.: inviting a colleague]
4. [flow 4, e.g.: search and filtering]

For each task:
- Scenario: realistic situation (2–3 sentences)
- Task: what the participant should do (no UI hints)
- Success criterion: how the moderator will know the task is complete
- Maximum time: how many minutes to allow

Rules:
- Tasks go from simple to complex
- First task is a warm-up (simple, to reduce anxiety)
- Wording describes the user's goal, not actions in the interface
- No UI terminology from the product (not "click the hamburger menu,"
  but "find the settings")
- Each task tests one specific flow
```

### Calibrating Task Difficulty

After generating tasks, check the difficulty balance. AI tends to skew everything toward one end.

```
Evaluate this set of usability test tasks against these criteria:

[insert generated tasks]

Evaluation criteria:
1. Is there a warm-up task (>90% participant success rate)?
2. Is there a medium-difficulty task (50–80% success)?
3. Is there a challenging task (<50% success)?
4. Do the wording choices contain hints (UI terminology)?
5. Do the tasks cover different parts of the product?
6. Is the task order realistic for a new user?

For each issue: describe what's wrong and suggest a revised wording.
```

## Think-Aloud Questions and Intermediate Probes

The think-aloud protocol generates the bulk of qualitative data. The participant voices their thoughts while completing a task. The moderator stays out of it — but sometimes needs to nudge the participant back to verbalizing.

A few phrases worth memorizing:

1. **When the participant goes silent:** "What are you looking for right now?" (not "What are you doing?" — that puts pressure on them)
2. **When they're struggling:** "What did you expect to see here?" (this captures the mental model)
3. **After they finish:** "Was that what you expected?" (catches hidden dissatisfaction)

After each task, 2–3 follow-up questions help. Generate them for the specific tasks you're running.

```
For each usability test task, generate 2–3 post-task questions.

Tasks:
[insert task list]

Question types:
- Subjective difficulty rating: "How easy or difficult was that?"
- Expectation match: "Did that work the way you expected?"
- Alternatives: "How else might you have tried to do that?"
- Confidence: "Are you confident the task is complete?"

Rules:
- Open-ended questions (no yes/no)
- Neutral wording (not "Did you like it?")
- No more than 3 questions per task (participants fatigue)
```

### Single Ease Question (SEQ) Scale

After each task, use the SEQ: "How easy was it to complete this task?" Scale of 1 to 7. Industry average is 5.5 ([MeasuringU data](https://measuringu.com/evolution-of-seq/)). Anything below that signals a problem.

SEQ gives your qualitative observations a quantitative anchor. "The participant gave a 3/7 and said they couldn't find the filters" carries more weight than "the participant couldn't find the filters."

## Closing Survey: The Full Picture

After all tasks, participants shift from doing to reflecting. This is where you gather data about the overall product experience.

```
Generate 6–8 questions for the closing survey of a usability test.

Product: [name]
What was tested: [list of flows]

Questions should cover:
- First impression of the product
- The most difficult moment during testing
- The clearest or most pleasant moment
- What the participant would change first
- Likelihood of recommendation (NPS or equivalent)
- Comparison to their current way of solving the problem

Format: open-ended questions.
Tone: conversational, not academic.
Final question: "Is there anything I didn't ask about that you'd like to share?"
```

The last question ("Is there anything I didn't ask?") regularly turns up insights the script missed entirely. Don't skip it.

## Prompt for a Full Script: Everything in One

Once you've got the structure down, you can generate the full script in one shot. This works best for experienced researchers who know exactly what to edit afterward.

```
You are a senior UX researcher with 10+ years of experience
moderating usability tests.

Create a complete usability test script for:
Product: [name and description, 2–3 sentences]
Product stage: [prototype / MVP / production]
Target audience: [description]
Key hypotheses to validate:
1. [hypothesis 1]
2. [hypothesis 2]
3. [hypothesis 3]

Testing format: [remote via Zoom / in-person]
Session duration: [45/60 minutes]

Script structure:
1. Introduction (verbatim text for the moderator)
2. Pre-test survey (5–7 open-ended questions)
3. 5–7 tasks (scenario + wording + success criterion + time)
4. Post-task questions (2–3 after each task + SEQ)
5. Closing survey (6–8 questions)
6. Wrap-up (verbatim text)

Additionally:
- Session timeline (minute by minute)
- Moderator checklist (preparation before the session)
- Observation sheet (table for notes during the test)

Rules:
- Tasks describe goals, not UI actions
- Questions are open-ended and neutral
- First task is a warm-up
- Order: simple to complex
- No product UI terminology in task wording
```

You get a 3–5 page document ready to use after 15–20 minutes of editing. That editing step isn't optional — the model doesn't know your product's specifics the way you do.

## Analyzing Usability Test Results with AI

Raw data needs structured analysis. One test generates 45–60 minutes of recording, dozens of observations, and subjective ratings. AI speeds up the move from raw data to things you can actually act on.

### Step 1: Transcription and Markup

Session transcription (Otter.ai, Fireflies, Zoom's built-in tools) gives you text. That text needs to be marked up by task.

```
Mark up this usability test transcription by block.

Transcription:
[insert text]

Test tasks:
[task list]

For each block, identify:
- Which task was being performed
- Start and end timestamp
- Status: completed / completed with difficulty / not completed
- Key participant quotes (verbatim)
- Moments of difficulty (what specifically caused the problem)
- Moments of success (what worked well)
```

### Step 2: Problem Matrix

After 3–5 sessions, a problem matrix starts to take shape. Each problem gets a severity and a frequency.

```
Based on usability test data, create a problem matrix.

Participant data:
[summary for each participant: tasks, statuses, issues, SEQ scores]

For each problem, define:
- Problem description (1–2 sentences)
- Severity: critical / serious / minor
  (critical = blocks task completion,
   serious = slows down but doesn't block,
   minor = annoying)
- Frequency: how many participants out of how many
- Tasks where the problem appeared
- Recommended fix

Sort by: severity first, then frequency.
Format: table.
```

The matrix answers the one question that matters: what do you fix first. A critical problem affecting 4 of 5 participants beats a minor problem affecting all 5.

### Step 3: Stakeholder Report

Test results need to reach the team. Twenty-page UX reports don't get read. What actually works: executive summary + top 5 problems + concrete recommendations.

```
Write an executive summary based on usability test results.

Product: [name]
Number of participants: [N]
Flows tested: [list]
Key metrics:
- Average Task Success Rate: [%]
- Average SEQ: [number]
- Average time on task: [min]

Problem matrix:
[insert matrix]

Report structure:
1. One-line verdict (product is/isn't ready for release)
2. Top 5 problems with recommendations (1–2 sentences each)
3. What's working well (2–3 points)
4. Recommended next steps (prioritized list)

Length: maximum 1 page.
Tone: factual, no value judgments.
Use numbers, not adjectives.
```

## Common Script Writing Mistakes

Even AI-generated scripts have predictable failure patterns. Check every script against this list before running it.

**Hints in task wording.** "Use the search bar to find a product" tells participants there's a search bar. Better: "Find wireless headphones under $50."

**Tasks that are too vague.** "Explore the app" tests nothing. There's no success criterion and no observable behavior to capture.

**Leading questions.** "Did you like the new design?" assumes it should be liked. Better: "Describe your impression of the interface."

**Too many tasks.** More than 7–8 in an hour wears participants down. The last tasks get less attention and the data drifts.

**No pilot run.** The first run of any script always surfaces wording problems. Run a pilot with a colleague before you bring in real participants.

## Moderator Checklist

How you prepare before the session matters as much as the script itself.

**Before the session:**
- Script printed or open on a second screen
- Prototype/product loaded, test account created
- Recording set up and tested
- Observation sheet ready (task, time, status, notes)
- Backup communication method (in case Zoom drops)
- Participant incentive prepared

**During the session:**
- Follow the script, but adjust question order if needed
- Don't help the participant — even when you really want to
- Log start and end time for each task
- Write down verbatim quotes, not your interpretation of what they meant
- Pause 3–5 seconds after a participant answers — a follow-up often comes

**After the session:**
- Write down your impressions within 15 minutes, while they're still fresh
- Flag the most memorable moments
- Update the observation sheet

## Integration with the Product Process

A usability test only creates value when results change decisions. Without that integration, it's just an expensive exercise.

Where testing actually fits into the process:

1. **Before development.** Testing [Figma prototypes](/blog/figma-to-prototype/). Fixing issues at this stage is cheap. Use [context engineering](/blog/context-engineering-guide/) to tailor prompts to the specific prototype.

2. **During the sprint.** Testing an MVP or staging version. Results go into the backlog as prioritized bugs.

3. **After release.** Testing the live product with real usage data. Validates — or challenges — the decisions made earlier.

A good cadence: one round (5 participants) every 2–4 weeks. That keeps data flowing without overwhelming the team.

## Metrics to Track

Quantitative metrics give you a baseline to compare across iterations.

| Metric | What It Measures | Benchmark |
|--------|-----------------|-----------|
| Task Success Rate | % of participants who completed the task | >78% (industry average) |
| Time on Task | Time to complete the task | Task-dependent |
| SEQ (Single Ease Question) | Subjective ease | >5.5 out of 7 |
| Error Rate | Number of errors per task | <0.7 per task |
| SUS (System Usability Scale) | Overall product usability | >68 out of 100 |

The [Task Success Rate (78%)](https://measuringu.com/task-completion/) and [SUS (68)](https://measuringu.com/sus/) benchmarks come from Jeff Sauro / MeasuringU research across hundreds of studies — treat them as a reference point, not a universal bar for your specific product.

Compare metrics across rounds. "Task Success Rate for registration went from 60% to 90% after the form redesign" says more than "we improved registration."

## Summary

AI produces a usability test script in 30–40 minutes instead of several hours. Script quality depends entirely on prompt quality: specific product, specific hypotheses, specific flows.

Steps:

1. Define 3–5 key flows to test
2. Generate tasks (no hints, simple to complex)
3. Check the script against the error list
4. Run a pilot with a colleague
5. Conduct 5 sessions
6. Analyze results with a problem matrix
7. Write an executive summary for the team

Five participants, one solid script, 45 minutes per session. That's enough to find 85% of the usability problems in your product.

---

*Need help setting up a usability testing process with AI? I help startups build AI products and automate processes — [belov.works](https://belov.works).*
