Hiring · August 4, 2026 · 10 min read
Mean Time to Hire: Why the Average Misleads You
Mean time to hire hides the reqs that ruin your quarter. Why the average misleads, how it diverges from the median, and where the days actually accumulate.
← Part of The five pillars of hiring: what assessments measure
On this page
- What is mean time to hire, and how do you calculate it?
- A worked example: five requisitions, two very different stories
- Why does the mean mislead so often?
- Where do the days in the tail actually accumulate?
- How do specialised tests compress the screening stage specifically?
- What compressing the screening stage will not fix
- Putting it together
If you are the hiring manager, talent leader or founder who has to report mean time to hire to the board, this post is about the trap hiding in that single number. The average looks reassuring precisely when it should not. You quote twenty-eight days, the room nods, and nobody can see that three requisitions sat open for eleven weeks each while the rest closed briskly. The mean smoothed them into invisibility — not a flaw in your arithmetic, but a flaw in reporting one figure where two are needed. This guide separates the mean from the median, shows with a worked example how far the two can drift apart, explains why the average misleads more often than it helps, and then gets specific about where the days that inflate it actually accumulate — and which of them a specialised assessment can give back.
A quick boundary. This post is about how you measure hiring speed, not what the terms mean at the requisition level; if you need the definitions and formulas, time to hire vs time to fill owns that ground. Here the argument is one level up: given a set of hires, which summary number should you trust, and what does the gap between the two tell you.
What is mean time to hire, and how do you calculate it?
Mean time to hire is the arithmetic average days your hires took, from a candidate entering the pipeline to accepting the offer. You take the duration of each completed hire, add them together, and divide by the number of hires. It is the default figure on most dashboards because it is trivial to compute and moves sensibly when the whole process speeds up or slows down.
The median is the other summary, and it answers a subtly different question. Sort every hire by how many days it took, from fastest to slowest, and the median is the value sitting in the middle — half your hires were quicker, half slower. Where the mean asks "what is the total divided evenly across everyone," the median asks "what did a typical hire actually feel like." On a tidy, symmetric set of durations the two land in almost the same place. Hiring data is rarely tidy or symmetric, which is the whole point.
A worked example: five requisitions, two very different stories
The cleanest way to feel the difference is to run the numbers on a small set of hires. The figures below are illustrative — invented to expose the arithmetic, not drawn from any survey or dataset. Imagine a team that closed five requisitions last quarter and recorded, for each, the days from application to offer acceptance.
- Requisition A — a customer support representative — closed in 19 days.
- Requisition B — an operations analyst — closed in 22 days.
- Requisition C — a data analyst — closed in 24 days.
- Requisition D — a niche machine learning engineer — closed in 71 days.
- Requisition E — a second machine learning engineer, backfilling a departure — closed in 74 days.
Now the two summaries. The median is the middle value once you sort them: 19, 22, 24, 71, 74 — the centre is 24 days. That is what a typical hire on this team felt like. The mean is the total, 210 days, divided by five: 42 days. Same five hires, two headline figures, eighteen days of daylight between them. Report 24 and you tell a story of a brisk process; report 42 and you tell a story of a sluggish one. Both are computed correctly from the same data, and neither, alone, is honest.
The divergence is not noise. It is information. Those two machine-learning searches took roughly three times as long as everything else, and the mean is the only summary that noticed; the median stepped right over them. Read together, the two numbers tell you exactly what happened: most of your hiring is fine, and two specific requisitions are in trouble. Read either alone and you lose half of that. The useful practice is never mean-or-median; it is mean-and-median, with the gap between them as its own diagnostic.
The median tells you what a typical hire felt like. The mean tells you whether a few reqs are quietly on fire. When the mean sits well above the median, that gap is not a rounding artefact — it is the tail of stuck requisitions, and it is the most actionable thing on your dashboard.
Why does the mean mislead so often?
The mean misleads because hiring durations are right-skewed by nature. Most hires cluster in a sensible band, but there is no ceiling on how long a single requisition can drag: a hard-to-source role, a candidate who takes three weeks to decide, an interviewer on leave, an approver who never quite replies. There is, however, a floor — no hire completes in negative days. That asymmetry produces a long right tail, and the mean is the one summary dragged toward it while the median holds its ground.
This matters more the smaller your sample. A team hiring five people a quarter can have its mean wrenched around by one unlucky search, which makes quarter-to-quarter mean comparisons treacherous — a jump from 28 to 42 days might be a broken process, or one niche req that will never recur. The median shrugs that off. There is also an incentive problem: the mean is the number most often reported upward, so it is the number most often gamed — close the easy reqs fast, let the hard ones sit, and the headline average improves while nothing underneath it has. The median is harder to flatter, and the gap between the two is nearly impossible to fake.
Do not compare a mean this quarter to a mean last quarter without also checking the median and the tail. On small samples, a single outlier requisition can swing the average by a week or more — enough to invent a crisis, or hide one, that has nothing to do with how your process actually performs.
None of this means you should abandon the mean. It responds to exactly the thing the median ignores, and the slow tail is where your real cost and candidate loss live. The discipline is to report the median as your headline, carry the mean alongside it, and treat any meaningful gap between them as an instruction: go and open up the requisitions in the tail. To compute both across your own dates while keeping the definition fixed, the time-to-hire calculator does the subtraction and the summaries for you.
Where do the days in the tail actually accumulate?
Once the mean has flagged a tail, the next question is where inside those stuck requisitions the days went. The instinct is to blame the interviews, because interviews are the visible, effortful part of hiring. They are almost never the culprit. Pull apart a requisition that took seventy days and the interviewing itself usually accounts for a handful of them. The rest is waiting.
Two stages dominate the tail. The first is the screening queue — applications sitting unread while someone finds the hours to work through them, plus the days spent hand-building a test relevant enough to trust for a niche role. That queue grows with the applicant pool. The second is calendar coordination: the req that sat eleven days waiting for a second interviewer's calendar to open, the offer that waited on an approver's inbox, the panel that could not align three diaries inside a fortnight. Neither is the interview. Both are dead time, and dead time is what stretches a tail.
- The screening queue — CVs waiting for a human to read them, and a role-relevant test waiting to be built by hand before anyone can be assessed.
- Calendar coordination — the gaps between stages while panels align diaries, an interviewer returns from leave, or an approver gets round to the sign-off.
- The interviews themselves — real time, but rarely the stage that produces a seventy-day outlier.
The distinction is worth holding onto because it determines what you should fix. Adding interviewers does nothing for a queue forming upstream of the interview. Rushing the interviews trades signal for a day or two that was never the problem. If the tail is made of screening delay and scheduling gaps — and it usually is — then those are the only two places worth attacking, and one of them is far more tractable than the other.
How do specialised tests compress the screening stage specifically?
The screening stage is the most compressible part of the tail because it is the part that is mechanical and serial rather than judgemental. Reading CVs one at a time and hand-building a test for a niche role are exactly the tasks that scale badly with volume and reward automation without costing you any decision quality. This is the stage a specialised, role-shaped assessment attacks.
Two things happen at once. First, the assessment no longer has to be built by hand: an AI-native skills assessment platform generates a job-relevant test per opening — the specialised assessment for that niche machine learning role exists in minutes, not days. Per-job generation turns the setup step from a gate into a non-event. Second, the whole pipeline is evaluated against that consistent assessment in parallel rather than one CV at a time, so the screening delay stops scaling with the applicant pool, and the queue that was dragging the tail stops forming. The human judgement stays where it earns its keep, on the shortlist. For which stages shrink and which must not, reduce time to hire with AI works through it stage by stage; the point specific to the mean is narrower — the compressible days are concentrated in the very stage that produces your slow tail.
The reason a specialised test matters here, rather than a generic aptitude quiz, is that a role-shaped assessment filters on the actual work. When the test mirrors the job, the shortlist it produces is one you can trust enough to move on quickly, which is the whole point of compressing the stage. A filter you do not trust just relocates the delay into the interviews, as the panel re-does the screening the assessment failed to do. You can watch a per-job assessment being generated and run end to end in a demo.
Attack the tail, not the average. Find the two or three requisitions inflating your mean, identify the stage each one stalled at, and fix that stage. Compressing the screening queue on your slowest reqs pulls the mean down far harder than any across-the-board push for a faster process, because that is where the excess days are actually sitting.
What compressing the screening stage will not fix
Honesty about the limits is what keeps this from becoming a sales pitch. A specialised assessment compresses one stage of the tail — the mechanical, screening part — and it is genuinely powerful there. It does nothing for the other things that stretch a requisition, and pretending otherwise would set you up to blame the wrong tool when the number does not move.
- Requisition approval delays. If a role sits for three weeks between headcount sign-off and going live, that is upstream of everything an assessment touches — it is a time-to-fill problem, not a screening one.
- Slow decision-making. A panel that takes a fortnight to agree on a candidate it has already assessed is a governance problem; better evidence helps the debate, but it cannot force a decision-maker to decide.
- Offer negotiation and notice periods. The days between a verbal yes and a signed contract, and the candidate's own notice period, are outside your process entirely.
- Scheduling gaps between the interviews you keep. Tightening the calendar is a coordination discipline; an assessment shortens the queue before the interviews, not the diary friction between them.
That list is the point, not a disclaimer. When you break your slow requisitions into stages and find the delay sitting in approval or negotiation, no amount of screening compression will help; the right move is to look upstream or at your decision governance instead. Separating the stages stops you from spending a screening fix on a scheduling problem. And none of this is a licence to chase a lower mean by cutting signal. The cost of a bad hire — the US Department of Labor puts it at least thirty per cent of first-year earnings, with other estimates at one to two times salary — dwarfs the few days you would save by skipping the assessment on your slow reqs. Compress the mechanical stage; never the judgement.
Putting it together
Report the median as your headline and carry the mean beside it, because the gap between them is the most useful thing on the dashboard. A mean sitting well above the median is not an error to reconcile; it is the tail of stuck requisitions announcing itself. Open those reqs up, find the stage each stalled at, and you will nearly always land on the screening queue or a scheduling gap rather than the interviews. The screening part is the compressible one — a specialised, per-job assessment makes the test exist in minutes and evaluates the pipeline in parallel, so the queue stops forming. Fix the tail and the mean falls honestly, on the reqs that were genuinely broken, without touching the healthy middle or the judgement that earns its place at the end.
A mean can hide three burning requisitions inside a comfortable-looking number. Measure both, read the gap, and go straight for the tail — because the average will never tell you which reqs are on fire, only that, on balance, the room can stay calm.
Written by
Aayesha Patel · Co-founder, Hanzomon Inc
Co-founder of Hanzomon. Writes about skills-based hiring, fair assessment and building a better candidate experience.