# Work sample test: design a fair, job-relevant hiring task

> A practical guide to work sample tests: job analysis, realistic scope, scoring rubrics, calibration, accommodations, AI rules, and bias monitoring.
> Author: Roman Belov · Published: 2026-06-17 · Source: https://futurecraft.pro/blog/work-sample-test/

A take-home assignment can show how someone works. It can also test stamina, spare time, familiarity with your private conventions, or willingness to do unpaid work.

Calling it a “work sample” does not settle which one you built.

A work sample worth defending starts with the job, uses a task small enough to respect the candidate's time, and records evidence before reviewers discuss impressions. The assignment, rubric, administration, and decision rule form one assessment. Fixing only the prompt is not enough.

## What the research does — and does not — say

Work samples have a solid research base, but the familiar “validity 0.54” claim is outdated. Sackett and colleagues revisited earlier personnel-selection meta-analyses and found that range-restriction corrections had often inflated estimates. Their revised operational validity estimate for work samples was **0.33**; structured interviews were **0.42** ([Sackett et al., 2022](https://doi.org/10.1037/apl0000994)).

Those are correlations aggregated across studies, not a guaranteed score for your assignment. Validity changes with the role, task, criterion, applicant population, and administration. The work-sample estimate also drew on research dominated by concurrent studies of existing employees, a limitation the paper discusses.

The practical conclusion is narrower than “work samples predict performance”:

> A job-relevant, consistently administered work sample can add useful evidence. Your company still has to show that its version measures what the decision requires.

Pair it with a [structured interview](/blog/structured-interview-kit/) rather than treating either method as a complete hiring system. The two can inspect different evidence: demonstrated task performance and past or situational behavior.

## Step 1: map the job before writing the assignment

Start with three inputs:

1. essential tasks performed in the role;
2. conditions under which those tasks happen;
3. observable competencies needed to do them well.

Interview people who do the work and their manager. Review recent incidents: launches, production failures, customer escalations, planning decisions. Ask what separated an adequate response from a harmful one.

This should use the same role evidence as the [job description](/blog/ai-job-descriptions/). If the vacancy promises strategic ownership but the test rewards ticket throughput, one of them describes the wrong job.

Then create a compact competency map:

| Essential task | Evidence to observe | Exclude from this test |
|---|---|---|
| Diagnose an API failure | prioritization, hypothesis quality, use of logs | memorized framework trivia |
| Propose a product change | evidence use, trade-offs, measurement plan | knowledge of internal acronyms |
| Edit a campaign brief | audience judgment, reasoning, clarity | visual design unless the role requires it |

The last column matters. A timed writing task can accidentally measure typing speed. An unfamiliar IDE can measure setup knowledge. A dense text-only brief can create an accessibility barrier unrelated to the job.

In the United States, the EEOC says selection procedures should be job-related and appropriate for their purpose; procedures with disparate impact require legal justification, and tests must accommodate qualified people with disabilities where required. Other jurisdictions have their own rules. Treat this article as design guidance, not legal advice, and review the process with local counsel or a qualified assessment professional.

## Step 2: choose the smallest representative slice

Do not begin with “we give everyone four hours.” Begin with the evidence you need.

A useful task should:

- resemble an important part of the actual role;
- produce observable choices, not just polished prose;
- include enough context to avoid testing guesswork;
- fit the stage of the hiring process;
- stop once the rubric has enough evidence.

Pilot completion time instead of declaring a universal limit. One debugging exercise may reveal useful evidence in 45 minutes. A portfolio review plus live discussion may be fairer than another take-home task. A complex leadership simulation may require more time and compensation.

Use fictional or sanitized inputs. Do not ask candidates to solve an unresolved company problem and then reuse their work. If the exercise is substantial, resembles billable work, or creates usable output, pay for it under a clear agreement. “We promise not to use it” does not return the candidate's evening.

Work samples are often viewed more favorably than many other selection methods, especially when candidates see the connection to the job. That does not excuse excessive scope; it makes visible relevance part of the design ([Hausknecht, Day & Thomas, 2004](https://doi.org/10.1111/j.1744-6570.2004.00003.x)).

## Step 3: standardize the candidate packet

The packet should make the task executable without rewarding candidates who know how to extract missing requirements from a recruiter.

Include:

- role and scenario context;
- the task and decision to make;
- input files and data dictionary;
- required deliverable and format;
- expected time budget;
- allowed tools and external resources;
- AI-use policy;
- submission and deletion process;
- accommodation contact;
- how the work will be evaluated.

Candidates need comparable conditions. That can mean the same task during one hiring cycle or equivalent forms drawn from a tested bank. Equivalent forms should target the same competencies, use the same rubric, and be piloted for similar difficulty. Rotating arbitrary assignments prevents comparison; reusing one public task forever rewards rehearsal.

Place the test after a narrow [resume screen](/blog/resume-screening-ai/) or eligibility check. Asking every applicant for unpaid work creates avoidable effort for candidates and reviewers without improving the evidence used at the first gate.

Do not change requirements after seeing a submission. If a clarification is necessary, send it to every active candidate and record the change.

## Step 4: build a rubric from observable evidence

A rubric is not a list of admirable nouns. “Strategic,” “senior,” and “good communicator” tell reviewers little.

Use a small set of competencies that came from the job analysis. For each one, describe evidence at distinct levels:

| Competency | Weak evidence | Meets the bar | Strong evidence |
|---|---|---|---|
| Diagnosis | Jumps to a fix without testing assumptions | Uses available evidence to isolate a plausible cause | Tests competing explanations and states remaining uncertainty |
| Trade-offs | Presents one answer as costless | Names material benefits, costs, and constraints | Compares viable options and identifies the reversible decision |
| Verification | Reports output without checking it | Checks the main path and important failure cases | Explains what was checked, what remains unverified, and why |

Decide before launch:

- which competencies are must-have;
- whether scores are weighted;
- what evidence triggers each rating;
- how missing or unusable work is handled;
- whether the result is advisory or a decision threshold;
- who can override the rule and how an override is documented.

Avoid decorative precision. A weighted score of 3.78 is not more scientific if reviewers cannot distinguish a 3 from a 4. The rubric makes judgment explicit and reviewable; it does not turn judgment into an objective measurement.

## Step 5: pilot the entire process

Run the candidate packet, not just the task, with several people who understand the role. Include different experience levels if the role allows them.

Record:

- actual completion time;
- instructions that required clarification;
- evidence each section produced;
- criteria that never affected a decision;
- accessibility or tooling friction;
- independent reviewer scores and reasons.

Have at least two reviewers score the same pilot submissions independently before they discuss them. Look at both numerical agreement and where their evidence notes differ. Averages can hide a serious disagreement: two reviewers may choose the same score for incompatible reasons.

Revise ambiguous anchors, irrelevant tasks, and misleading inputs. Do not require a current employee to receive an arbitrary minimum score; the pilot is a design test, not proof that every incumbent fits a new scale.

## Step 6: state the AI rule as part of the construct

“AI allowed” is too vague. Define the environment you intend to measure.

If the job uses AI, you might permit a named set of tools and require candidates to submit:

- which tools they used;
- the important prompts or workflow;
- changes made to generated output;
- checks performed;
- unresolved uncertainty.

Then score the capability the role needs: framing, verification, correction, security judgment, and ownership of the final result. Do not score prompt length or fluent AI prose as competence.

If the goal is an unaided baseline, explain why AI is prohibited and which aids are still permitted. Apply the same rule to every candidate. Do not infer tool use from writing style or feed submissions into an AI detector; neither is a reliable substitute for declared conditions and evidence.

AI can draft the assessment, but it must not invent job requirements:

```text
You are drafting a work sample from supplied evidence.

Inputs:
- essential tasks and critical incidents
- competency map
- approved time and tool constraints
- examples of acceptable and harmful outcomes

Produce:
1. one representative task;
2. a candidate packet with all required inputs;
3. a rubric using observable behavioral anchors;
4. two equivalent task variants;
5. risks: construct-irrelevant difficulty, accessibility, data privacy,
   unpaid production work, and likely information leakage;
6. questions that a human owner must resolve.

Rules:
- Do not invent company facts, benchmarks, legal conclusions, or score cutoffs.
- Mark missing information as UNKNOWN.
- Keep every criterion traceable to an essential task.
```

A hiring manager and someone familiar with employment law, accessibility, or assessment design should review the result before candidates see it.

## Step 7: administer and score without improvising

Give every candidate the same instructions, deadline logic, tool policy, and access to clarification. Offer a clear route to request an accommodation. An accommodation should remove an irrelevant barrier, not lower the job-related standard.

For scoring:

1. remove identifying information where feasible and lawful;
2. assign reviewers before opening submissions;
3. have reviewers record evidence and score independently;
4. reconcile disagreements only after independent scoring;
5. document the final decision against the rubric;
6. retain and delete data according to a declared policy.

Blind review can reduce exposure to some identity cues, but it is not a complete bias control. Portfolios, writing, and context can still reveal identity; removing context can also erase evidence relevant to the work. Use it where it improves the process, not as a ritual.

Keep the follow-up discussion structured. Ask candidates to explain choices, correct an error, or respond to a changed constraint. Do not turn it into an unscored interrogation that silently overrides the rubric.

## Step 8: validate the test after launch

A pilot checks whether the mechanism runs. Operational data tells you whether it helps.

Track by task version and hiring stage:

- invitation and completion rates;
- candidate withdrawals and stated reasons;
- reviewer agreement;
- score distributions;
- pass rates and later hiring decisions;
- performance evidence after hire, when lawful and meaningful;
- adverse impact across protected groups, using appropriately governed data;
- candidate complaints and accommodation failures.

Small samples create noisy percentages. Do not celebrate or condemn a test after three hires. Preserve raw counts, review patterns over time, and get specialist help before interpreting sensitive demographic data.

If a criterion never predicts a decision, reviewers interpret it differently, or it disproportionately excludes a group without sufficient job-related justification, investigate it. The answer may be a better anchor, a different task, another selection method, or retirement of the test.

US employers remain responsible for selection tools supplied by vendors, including AI systems. The EEOC's guidance warns that automated tools can create disability barriers or adverse impact; a vendor's accuracy claim does not replace the employer's review.

## A compact example: backend incident triage

**Scenario:** A service's error rate increased after a release. The candidate receives sanitized logs, a deployment diff, a simple architecture diagram, and three monitoring charts.

**Task:** Produce a one-page incident note with the leading hypotheses, next diagnostic actions, an immediate risk-control decision, and unknowns. Then discuss it for 20 minutes with two reviewers.

**What it measures:** evidence use, prioritization, trade-offs, communication, and verification.

**What it does not measure:** production access speed, familiarity with the company's internal tools, or ability to implement a full fix.

**AI policy:** use is permitted and disclosed. Reviewers assess whether the candidate verifies generated suggestions against the supplied evidence.

This version does not ask for a deployable feature. It still gives reviewers concrete decisions to inspect.

## Pre-launch checklist

- [ ] Every criterion traces to an essential task.
- [ ] The candidate packet contains the context needed to perform.
- [ ] Scope was timed in a pilot and unnecessary work was removed.
- [ ] Inputs are fictional, licensed, or sanitized.
- [ ] Substantial or usable work has a compensation policy.
- [ ] AI and tool rules are explicit and consistent.
- [ ] An accommodation route is visible.
- [ ] Equivalent variants use the same competency map and rubric.
- [ ] Reviewers practiced independent evidence-based scoring.
- [ ] Data retention, access, and deletion are defined.
- [ ] Outcome and adverse-impact review has an owner.
- [ ] Local legal and accessibility requirements were checked.

A strong work sample does not imitate an entire working week. It captures a narrow decision that matters, gives each candidate a fair chance to show the relevant skill, and leaves an evidence trail you can audit.

## Sources

- [Sackett et al.: revised validity estimates for personnel selection](https://doi.org/10.1037/apl0000994)
- [Sackett et al.: implications for selection-system design](https://doi.org/10.1017/iop.2023.24)
- [Roth et al.: meta-analysis of work-sample test validity](https://doi.org/10.1111/j.1744-6570.2005.00672.x)
- [Hausknecht, Day & Thomas: meta-analysis of applicant reactions](https://doi.org/10.1111/j.1744-6570.2004.00003.x)
- [EEOC: employment tests and selection procedures](https://www.eeoc.gov/laws/guidance/employment-tests-and-selection-procedures)
- [EEOC: reasonable accommodation in selection](https://www.eeoc.gov/laws/guidance/enforcement-guidance-reasonable-accommodation-and-undue-hardship-under-ada)
- [EEOC: AI and the ADA](https://www.eeoc.gov/eeoc-disability-related-resources/artificial-intelligence-and-ada)
