Work sample test: design a fair, job-relevant hiring task

By Updated

What is a work sample test?

A work sample test asks a candidate to perform a small, representative part of the job under defined conditions. It is useful only when the task follows from real job requirements, candidates receive comparable instructions, and trained reviewers score observable evidence against a tested rubric.

TL;DR

  • -Start with the job, not a clever assignment. Map essential tasks and competencies before writing the test.
  • -Use the smallest realistic task that reveals enough evidence. There is no universal two-hour or four-hour rule.
  • -Give candidates equivalent inputs, constraints, tool access, and an explicit AI policy. Offer reasonable accommodations under applicable law.
  • -A rubric reduces discretion; it does not make judgment objective. Pilot the task and have reviewers score samples independently.
  • -Use fictional or sanitized inputs. Pay for substantial work and never use a candidate's output as free production work.
  • -Track completion, score agreement, hiring outcomes, and adverse impact. Revise or retire a test that does not work.

A take-home assignment can show how someone works. It can also test stamina, spare time, familiarity with your private conventions, or willingness to do unpaid work.

Calling it a “work sample” does not settle which one you built.

A work sample worth defending starts with the job, uses a task small enough to respect the candidate’s time, and records evidence before reviewers discuss impressions. The assignment, rubric, administration, and decision rule form one assessment. Fixing only the prompt is not enough.

What the research does — and does not — say

Work samples have a solid research base, but the familiar “validity 0.54” claim is outdated. Sackett and colleagues revisited earlier personnel-selection meta-analyses and found that range-restriction corrections had often inflated estimates. Their revised operational validity estimate for work samples was 0.33; structured interviews were 0.42 (Sackett et al., 2022).

Those are correlations aggregated across studies, not a guaranteed score for your assignment. Validity changes with the role, task, criterion, applicant population, and administration. The work-sample estimate also drew on research dominated by concurrent studies of existing employees, a limitation the paper discusses.

The practical conclusion is narrower than “work samples predict performance”:

A job-relevant, consistently administered work sample can add useful evidence. Your company still has to show that its version measures what the decision requires.

Pair it with a structured interview rather than treating either method as a complete hiring system. The two can inspect different evidence: demonstrated task performance and past or situational behavior.

Step 1: map the job before writing the assignment

Start with three inputs:

  1. essential tasks performed in the role;
  2. conditions under which those tasks happen;
  3. observable competencies needed to do them well.

Interview people who do the work and their manager. Review recent incidents: launches, production failures, customer escalations, planning decisions. Ask what separated an adequate response from a harmful one.

This should use the same role evidence as the job description. If the vacancy promises strategic ownership but the test rewards ticket throughput, one of them describes the wrong job.

Then create a compact competency map:

Essential taskEvidence to observeExclude from this test
Diagnose an API failureprioritization, hypothesis quality, use of logsmemorized framework trivia
Propose a product changeevidence use, trade-offs, measurement planknowledge of internal acronyms
Edit a campaign briefaudience judgment, reasoning, clarityvisual design unless the role requires it

The last column matters. A timed writing task can accidentally measure typing speed. An unfamiliar IDE can measure setup knowledge. A dense text-only brief can create an accessibility barrier unrelated to the job.

In the United States, the EEOC says selection procedures should be job-related and appropriate for their purpose; procedures with disparate impact require legal justification, and tests must accommodate qualified people with disabilities where required. Other jurisdictions have their own rules. Treat this article as design guidance, not legal advice, and review the process with local counsel or a qualified assessment professional.

Step 2: choose the smallest representative slice

Do not begin with “we give everyone four hours.” Begin with the evidence you need.

A useful task should:

  • resemble an important part of the actual role;
  • produce observable choices, not just polished prose;
  • include enough context to avoid testing guesswork;
  • fit the stage of the hiring process;
  • stop once the rubric has enough evidence.

Pilot completion time instead of declaring a universal limit. One debugging exercise may reveal useful evidence in 45 minutes. A portfolio review plus live discussion may be fairer than another take-home task. A complex leadership simulation may require more time and compensation.

Use fictional or sanitized inputs. Do not ask candidates to solve an unresolved company problem and then reuse their work. If the exercise is substantial, resembles billable work, or creates usable output, pay for it under a clear agreement. “We promise not to use it” does not return the candidate’s evening.

Work samples are often viewed more favorably than many other selection methods, especially when candidates see the connection to the job. That does not excuse excessive scope; it makes visible relevance part of the design (Hausknecht, Day & Thomas, 2004).

Step 3: standardize the candidate packet

The packet should make the task executable without rewarding candidates who know how to extract missing requirements from a recruiter.

Include:

  • role and scenario context;
  • the task and decision to make;
  • input files and data dictionary;
  • required deliverable and format;
  • expected time budget;
  • allowed tools and external resources;
  • AI-use policy;
  • submission and deletion process;
  • accommodation contact;
  • how the work will be evaluated.

Candidates need comparable conditions. That can mean the same task during one hiring cycle or equivalent forms drawn from a tested bank. Equivalent forms should target the same competencies, use the same rubric, and be piloted for similar difficulty. Rotating arbitrary assignments prevents comparison; reusing one public task forever rewards rehearsal.

Place the test after a narrow resume screen or eligibility check. Asking every applicant for unpaid work creates avoidable effort for candidates and reviewers without improving the evidence used at the first gate.

Do not change requirements after seeing a submission. If a clarification is necessary, send it to every active candidate and record the change.

Step 4: build a rubric from observable evidence

A rubric is not a list of admirable nouns. “Strategic,” “senior,” and “good communicator” tell reviewers little.

Use a small set of competencies that came from the job analysis. For each one, describe evidence at distinct levels:

CompetencyWeak evidenceMeets the barStrong evidence
DiagnosisJumps to a fix without testing assumptionsUses available evidence to isolate a plausible causeTests competing explanations and states remaining uncertainty
Trade-offsPresents one answer as costlessNames material benefits, costs, and constraintsCompares viable options and identifies the reversible decision
VerificationReports output without checking itChecks the main path and important failure casesExplains what was checked, what remains unverified, and why

Decide before launch:

  • which competencies are must-have;
  • whether scores are weighted;
  • what evidence triggers each rating;
  • how missing or unusable work is handled;
  • whether the result is advisory or a decision threshold;
  • who can override the rule and how an override is documented.

Avoid decorative precision. A weighted score of 3.78 is not more scientific if reviewers cannot distinguish a 3 from a 4. The rubric makes judgment explicit and reviewable; it does not turn judgment into an objective measurement.

Step 5: pilot the entire process

Run the candidate packet, not just the task, with several people who understand the role. Include different experience levels if the role allows them.

Record:

  • actual completion time;
  • instructions that required clarification;
  • evidence each section produced;
  • criteria that never affected a decision;
  • accessibility or tooling friction;
  • independent reviewer scores and reasons.

Have at least two reviewers score the same pilot submissions independently before they discuss them. Look at both numerical agreement and where their evidence notes differ. Averages can hide a serious disagreement: two reviewers may choose the same score for incompatible reasons.

Revise ambiguous anchors, irrelevant tasks, and misleading inputs. Do not require a current employee to receive an arbitrary minimum score; the pilot is a design test, not proof that every incumbent fits a new scale.

Step 6: state the AI rule as part of the construct

“AI allowed” is too vague. Define the environment you intend to measure.

If the job uses AI, you might permit a named set of tools and require candidates to submit:

  • which tools they used;
  • the important prompts or workflow;
  • changes made to generated output;
  • checks performed;
  • unresolved uncertainty.

Then score the capability the role needs: framing, verification, correction, security judgment, and ownership of the final result. Do not score prompt length or fluent AI prose as competence.

If the goal is an unaided baseline, explain why AI is prohibited and which aids are still permitted. Apply the same rule to every candidate. Do not infer tool use from writing style or feed submissions into an AI detector; neither is a reliable substitute for declared conditions and evidence.

AI can draft the assessment, but it must not invent job requirements:

You are drafting a work sample from supplied evidence.

Inputs:
- essential tasks and critical incidents
- competency map
- approved time and tool constraints
- examples of acceptable and harmful outcomes

Produce:
1. one representative task;
2. a candidate packet with all required inputs;
3. a rubric using observable behavioral anchors;
4. two equivalent task variants;
5. risks: construct-irrelevant difficulty, accessibility, data privacy,
   unpaid production work, and likely information leakage;
6. questions that a human owner must resolve.

Rules:
- Do not invent company facts, benchmarks, legal conclusions, or score cutoffs.
- Mark missing information as UNKNOWN.
- Keep every criterion traceable to an essential task.

A hiring manager and someone familiar with employment law, accessibility, or assessment design should review the result before candidates see it.

Step 7: administer and score without improvising

Give every candidate the same instructions, deadline logic, tool policy, and access to clarification. Offer a clear route to request an accommodation. An accommodation should remove an irrelevant barrier, not lower the job-related standard.

For scoring:

  1. remove identifying information where feasible and lawful;
  2. assign reviewers before opening submissions;
  3. have reviewers record evidence and score independently;
  4. reconcile disagreements only after independent scoring;
  5. document the final decision against the rubric;
  6. retain and delete data according to a declared policy.

Blind review can reduce exposure to some identity cues, but it is not a complete bias control. Portfolios, writing, and context can still reveal identity; removing context can also erase evidence relevant to the work. Use it where it improves the process, not as a ritual.

Keep the follow-up discussion structured. Ask candidates to explain choices, correct an error, or respond to a changed constraint. Do not turn it into an unscored interrogation that silently overrides the rubric.

Step 8: validate the test after launch

A pilot checks whether the mechanism runs. Operational data tells you whether it helps.

Track by task version and hiring stage:

  • invitation and completion rates;
  • candidate withdrawals and stated reasons;
  • reviewer agreement;
  • score distributions;
  • pass rates and later hiring decisions;
  • performance evidence after hire, when lawful and meaningful;
  • adverse impact across protected groups, using appropriately governed data;
  • candidate complaints and accommodation failures.

Small samples create noisy percentages. Do not celebrate or condemn a test after three hires. Preserve raw counts, review patterns over time, and get specialist help before interpreting sensitive demographic data.

If a criterion never predicts a decision, reviewers interpret it differently, or it disproportionately excludes a group without sufficient job-related justification, investigate it. The answer may be a better anchor, a different task, another selection method, or retirement of the test.

US employers remain responsible for selection tools supplied by vendors, including AI systems. The EEOC’s guidance warns that automated tools can create disability barriers or adverse impact; a vendor’s accuracy claim does not replace the employer’s review.

A compact example: backend incident triage

Scenario: A service’s error rate increased after a release. The candidate receives sanitized logs, a deployment diff, a simple architecture diagram, and three monitoring charts.

Task: Produce a one-page incident note with the leading hypotheses, next diagnostic actions, an immediate risk-control decision, and unknowns. Then discuss it for 20 minutes with two reviewers.

What it measures: evidence use, prioritization, trade-offs, communication, and verification.

What it does not measure: production access speed, familiarity with the company’s internal tools, or ability to implement a full fix.

AI policy: use is permitted and disclosed. Reviewers assess whether the candidate verifies generated suggestions against the supplied evidence.

This version does not ask for a deployable feature. It still gives reviewers concrete decisions to inspect.

Pre-launch checklist

  • Every criterion traces to an essential task.
  • The candidate packet contains the context needed to perform.
  • Scope was timed in a pilot and unnecessary work was removed.
  • Inputs are fictional, licensed, or sanitized.
  • Substantial or usable work has a compensation policy.
  • AI and tool rules are explicit and consistent.
  • An accommodation route is visible.
  • Equivalent variants use the same competency map and rubric.
  • Reviewers practiced independent evidence-based scoring.
  • Data retention, access, and deletion are defined.
  • Outcome and adverse-impact review has an owner.
  • Local legal and accessibility requirements were checked.

A strong work sample does not imitate an entire working week. It captures a narrow decision that matters, gives each candidate a fair chance to show the relevant skill, and leaves an evidence trail you can audit.

Sources

Frequently Asked Questions

How long should a work sample test take?
Use the shortest task that produces enough evidence for the hiring decision. Pilot it with people familiar with the role, measure actual completion time, and remove work that does not affect the rubric. If the task requires substantial effort or could benefit the company, pay candidates.
Should candidates be allowed to use AI?
Match the rule to the job. If the role uses AI, permit it under the same conditions for every candidate and assess prompting, verification, correction, and final judgment. If you need an unaided baseline, explain what is prohibited and why.
Should every candidate receive the same assignment?
Candidates need comparable conditions, not necessarily one permanently reused file. Equivalent forms can reduce leaks if they measure the same competencies, use the same rubric, and have been piloted for similar difficulty.
Does a scoring rubric remove hiring bias?
No. A rubric makes criteria explicit and decisions auditable, but biased criteria, ambiguous anchors, inaccessible administration, or inconsistent reviewers can still produce unfair results. Monitor outcomes and seek qualified legal or assessment advice where needed.