All posts

Hiring · July 20, 2026 · 10 min read

Work Sample Tests: The Most Predictive Way to Hire

Work sample tests measure what a candidate can actually do and predict job performance better than résumés or interviews. How to design and score them well.

By Aayesha Patel · Co-founder, Hanzomon Inc

Share

Part of The five pillars of hiring: what assessments measure

Hiring
On this page

If you could keep only one signal in your hiring process, the evidence says it should be a sample of the real work. For hiring managers and talent leaders under pressure to make faster, better decisions, that matters enormously: a work sample test puts candidates in front of a small, realistic slice of the job and scores what they produce. It is the difference between asking someone whether they can do the work and watching them do it — and it is why work sample tests consistently outrank résumés and unstructured interviews as predictors of who will actually perform.

Why work samples predict performance

The logic is almost too simple to trust: the best predictor of how someone will perform a task is how they perform that task. Decades of research into selection methods have ranked work samples and structured, job-relevant assessments among the strongest predictors of performance — well ahead of years of experience or education on a résumé. They also travel well across backgrounds, because they measure demonstrated ability rather than pedigree, which is one reason they help reduce bias in hiring rather than encode it.

Contrast that with the signals most processes lean on. A résumé tells you what someone says they have done. An unstructured interview tells you how well they interview — a skill that correlates only loosely with the job. A work sample cuts through both by generating direct evidence: here is a task, here is what the candidate produced, here is how it scored. That evidence is what makes the rest of the process sharper. This is the heart of skills-based hiring — deciding on demonstrated ability rather than proxies for it.

There is a fairness dividend, too, and it is worth stating plainly. Proxies like where someone went to university or which companies appear on their CV tend to track opportunity as much as ability, which quietly narrows a pipeline before anyone has done any actual work. A well-designed work sample resets that: two candidates from very different backgrounds face the same task and are judged on the same output. It will not fix every source of bias on its own, but replacing a pedigree signal with a performance signal is one of the most effective single changes a team can make to widen a funnel without lowering the bar.

A work sample does not have to mean a giant take-home. A focused, role-tuned assessment gathers the same signal in under an hour — see a sample assessment end to end before you design your own.

What makes a work sample good

Not every task that looks like real work is a good work sample. The ones that produce clean, defensible signal share four traits, and it is worth being deliberate about each before you send anything to a candidate.

  • Job-relevant. The task mirrors something the person will actually do — not a puzzle chosen because it is easy to grade. Decide up front what the role genuinely requires.
  • Right level. An entry screen and a senior loop should not look the same; difficulty and emphasis should track the seniority you are hiring for.
  • Realistic tools. Let candidates use what they would use on the job, including AI assistants, and assess how well they use them — see AI fluency as a hiring signal.
  • Consistently scored. A shared rubric across the five pillars of hiring keeps candidates comparable and the decision defensible.

That last point does the heaviest lifting. A work sample scored on gut feel is just an interview with extra steps. A work sample scored against a rubric — the same criteria applied to every candidate for the role — is where the predictive power and the fairness both come from. Standardisation is not bureaucracy; it is what makes the comparison mean anything.

Building the rubric is where most of the thinking should go, and it pays to do it before you see a single submission. Write down, in advance, what a strong answer looks like, what a passable one looks like, and what should worry you — then score against those criteria rather than reacting to whichever submission you happened to open first. Anchoring the rubric before scoring guards against the drift where standards quietly rise or fall depending on the order candidates arrive in. If more than one person is scoring, calibrate on a couple of real submissions together so you are applying the same bar rather than each grading to your own private taste.

Score capabilities, not vibes

The most common rubric mistake is scoring the artefact instead of the capability behind it. A single overall mark tells you a candidate did well without telling you why, which makes it almost impossible to compare two people who scored the same for different reasons. The fix is to name the handful of capabilities the task is genuinely probing and score each one separately — for an engineering sample, perhaps correctness, code clarity, handling of edge cases, and the quality of the trade-offs the candidate explained. Each gets its own short scale with concrete anchors — what a 1 looks like, what a 3, what a 5 — described in observable behaviour rather than adjectives like 'strong' that mean different things to different scorers. Keep the count small: three to five capabilities is usually plenty, because a rubric with fifteen dimensions spreads attention too thin to score any of them well.

Candidate score report showing a work sample assessment scored against a rubric
A scored work sample: the candidate's output measured against a consistent, role-specific rubric.

Take-homes, live exercises and completion

The format you choose shapes who finishes. Long unpaid take-homes gather rich signal but quietly bias your funnel toward candidates with the most free time — often the least representative slice of your applicant pool. Live exercises are fairer on time but can favour people who perform well under observation. There is no perfect format; there is a fit between the role, the seniority and the effort you are asking for. If you are weighing the trade-off directly, take-home assignments versus live coding lays out where each earns its place.

A useful rule of thumb: match the effort you ask for to the stage of the funnel. Early on, when you are working with a large pool and low commitment on both sides, keep the sample short and sharp — enough to separate signal from noise, no more. Reserve the deeper, more time-consuming exercise for later stages where the pool is small and both sides have already invested. Asking every applicant for a multi-hour assignment at the top of the funnel is the fastest way to lose the very people you most want, who tend to have the most competing options and the least patience for busywork.

Right-size the time by seniority

Seniority changes the maths on how much time you can reasonably ask for. Junior and graduate candidates are usually the most willing to invest in a longer exercise, because a sample is often their best chance to show ability that a thin CV cannot; a well-scoped task of around an hour sits comfortably here. As seniority rises, the calculus inverts. Senior candidates almost always have current jobs and a strong sense of their own market value; a two-hour unpaid assignment reads not as diligence but as a signal that you do not respect their time, and your best prospects will simply decline. For those roles, keep the sample short and high-signal, or move the deeper assessment into a paid or live format.

That raises the paid-versus-unpaid question, which turns almost entirely on length. A short, self-contained task that produces nothing you would ever ship is fair to run unpaid — it is a genuine assessment, not free labour. The line gets crossed when the exercise starts to look like real deliverable work; paying for that time removes the quiet bias unpaid work introduces towards people who can afford to give away an afternoon.

A work sample that takes six hours is not a stronger signal — it is a completion problem. If strong candidates drop out before finishing, you are selecting for availability, not ability. Keep the task tight and respect the candidate's time.

Work samples and interviews are partners

A work sample is not a replacement for human judgement — it is what focuses it. Use the sample to gather evidence and decide who to advance, then use a structured interview to probe the reasoning behind what they produced. The assessment tells you what a candidate can do; the conversation tells you how they think and whether they will thrive on your team. Run in that order, each round makes the next one better rather than repeating it.

This partnership also protects you from the two classic failure modes. Relying only on the sample can miss context — why a candidate made a particular call, or how they would work with a team. Relying only on the interview lets charisma stand in for competence. Together, they give you both the evidence and the story behind it, which is exactly what a defensible hiring decision needs.

Handling candidate objections

Even a well-designed sample will draw pushback, and how you respond says as much to a candidate as the task itself. The objections you hear most often have good answers.

  • 'This is unpaid work.' Fair when the task is long. Keep it short and clearly artificial, or pay for the deeper version — and say so up front.
  • 'I don't have time.' Publish the realistic time budget before they start and hold the task to it, so they can plan around a tightly scoped hour.
  • 'Why not just judge my portfolio?' Prior work is uncontrolled — you cannot see the brief, the help, or the timeline. A standardised sample makes candidates comparable on the same terms.
  • 'How will this be scored?' Tell them. Sharing the capabilities you are assessing mirrors real work, where good briefs state what success looks like.

The through-line is transparency. Explain why the task exists and how it fits the wider process, and most reasonable objections soften. If someone still refuses a short, well-scoped, clearly job-relevant task, that is itself a quiet signal worth noting rather than a failure of the method.

Reading the results without over-reading them

A work sample gives you a strong signal, not a verdict. Treat a single score as the whole truth and you will occasionally reject someone who had a bad hour and advance someone who had a lucky one. The safer read is to use the sample to sort candidates into confident advances, confident declines, and a middle band that the interview exists to resolve. Pay attention to how a candidate arrived at their answer, not just whether it was right — the approach often generalises to the job better than the specific output does. And be honest about what the task did and did not measure: a sample that tests writing tells you little about whether someone can run a meeting.

One task rarely covers a whole role. A work sample is strongest as one input among a small number of complementary signals — pair it with a structured interview and, where relevant, a reference on how the person works in a team. Breadth of evidence beats depth on any single dimension.

Making work samples fair and defensible

Because work samples generate a clear record of how each candidate was assessed, they are among the most defensible methods you can use — provided you standardise them. Same task, same conditions, same rubric, applied to every candidate for the role. That consistency is what lets you show your reasoning if a decision is ever challenged, and it is why well-run samples tend to widen rather than narrow your pool. When the task is genuinely job-relevant and scored the same way for everyone, you are measuring ability, not accent, pedigree or interview polish.

This defensibility is not only a legal comfort — it is an organisational one. When a hiring manager can point to what a candidate actually produced and how it was scored, disagreements about who to advance become conversations about evidence rather than clashes of opinion. That lowers the temperature of the whole process and makes it far easier to bring a sceptical stakeholder along. The record also compounds over time: as you run more samples, you learn which tasks predict well for which roles, and you can tighten them accordingly instead of relying on instinct that never gets tested.

Done well, work samples raise the quality of every downstream decision and lower your exposure to the cost of a bad hire. They shorten the path to a confident yes or no, because the evidence is in front of you rather than inferred. You can watch an assessment reshape by role and seniority to see how a work sample adapts in practice — and how the same rigour scales across a whole pipeline.

Stop asking whether candidates can do the job. Give them a small piece of it and watch.
Work sample testsSkills assessmentPredictive validityStructured hiring
A

Written by

Aayesha Patel · Co-founder, Hanzomon Inc

Co-founder of Hanzomon. Writes about skills-based hiring, fair assessment and building a better candidate experience.

Put this into practice

The assessments, role guides and calculators that turn what you have just read into a hiring decision.

Frequently asked questions

What is a work sample test?

A work sample test asks a candidate to perform a small, realistic piece of the actual job — writing a query, debugging code, drafting a customer reply, prioritising a backlog — under standardised conditions, and scores the result against a rubric. It measures what someone can do rather than what they claim on a résumé. Because it mirrors real tasks, it tends to predict on-the-job performance more reliably than experience or interview impressions alone.

Are work sample tests better than interviews?

For predicting job performance, samples of the real work are among the strongest single signals available, and they are less prone to interviewer bias than unstructured conversation. But it is not either-or. The best process usually combines a work sample to gather evidence with a structured interview to probe the reasoning behind it. The sample tells you what a candidate can do; the conversation tells you how they think.

How long should a work sample test be?

Long enough to gather real signal, short enough to respect candidates' time — often a focused 45 to 75 minutes. Long unpaid take-homes hurt completion rates and bias the funnel toward candidates with the most free time. A tightly scoped, role-tuned task gathers comparable signal in under an hour and keeps strong applicants engaged rather than dropping out midway.

Should candidates be allowed to use AI in a work sample?

In most modern roles, yes — because they will use AI on the job. Blocking it measures a skill the work no longer rewards. A better approach lets candidates use the tools they would normally reach for and assesses how well they use them: prompt craft, verification and judgement. If AI collaboration is central to the role, a dedicated live exercise gives you a cleaner read than an AI-banned test.

How do you make work sample tests fair?

Fairness comes from standardisation. Give every candidate for a role the same task under the same conditions, score against a shared rubric rather than gut feel, and make sure the task is genuinely job-relevant rather than a puzzle that favours a particular background. Consistent scoring keeps candidates comparable, reduces bias, and makes the eventual decision defensible if it is ever questioned.

Related posts

See it on your own job description

Join the early-access waitlist and watch H-Evaluate build an assessment for a real role.

See it on your own job description