Hiring · July 30, 2026 · 10 min read
Behavioural assessment and bias: what structure fixes
Unstructured behavioural interviews are where similarity bias and the halo effect live. How anchored scales and the same probes for everyone narrow that channel.
← Part of The five pillars of hiring: what assessments measure
On this page
- Where bias actually lives in behavioural interviews
- What structure narrows — and how
- The same probes for everyone
- Anchored scales, not adjectives
- Observed behaviour over claimed behaviour
- What structure does not fix
- Self-report can be gamed
- Anchors can encode one working style
- Verify the result, don't assume it
- Behavioural, in the context of the other pillars
Ask a hiring manager why they rated a candidate highly on 'teamwork' and often the honest answer is: I liked them. This post is for hiring managers and talent leaders who use behavioural interviews — the 'tell me about a time you…' questions meant to predict how someone works on a team — and want them to reduce bias rather than launder it. Behavioural assessment is where the softest judgements in hiring get made, which makes it the pillar where familiarity most easily passes for evidence. The claim worth defending is narrow and honest: structured behavioural assessment reduces bias relative to an unstructured interview, because it gives bias fewer places to operate. It does not eliminate it, and this post is as much about what structure cannot fix as what it can.
Why the stakes are high: behaviour drives tenure and performance as much as raw skill, yet it is the pillar most teams reduce to a gut feel formed in the first five minutes. When that gut feel is really a similarity signal, it quietly skews who advances — and because it wears the language of assessment, no one calls it bias. The fix is not to stop measuring behaviour. It is to measure it with structure, then check the outcome. This is a deep companion to our guide on how to reduce bias in hiring, focused on the behavioural pillar specifically.
Where bias actually lives in behavioural interviews
The unstructured behavioural interview is a hospitable environment for two well-documented biases. The first is similarity bias — the ordinary pull toward candidates who remind us of ourselves. A shared university, a familiar way of talking about work, a background that maps onto our own registers as 'a safe pair of hands' before a single answer has been weighed. In a free-flowing conversation, rapport builds fastest with the candidate you already resemble, and rapport is easy to mistake for a strong signal on collaboration or judgement.
The second is the halo effect. One vivid impression — a confident opening, an impressive former employer, a story told well — spreads a glow across every later judgement. A candidate who lands their first answer gets the benefit of the doubt on the next four. The problem is not that the impression is wrong; it is that it contaminates unrelated ratings. You think you are scoring five distinct behaviours, but you are really scoring the first one five times.
Both biases share a mechanism: an impression arrives before the evidence, and then the evidence gets read to fit the impression. That is 'I liked them' dressed as assessment. The interviewer is usually certain they were fair, which is exactly why willpower is not the fix. Awareness of a bias does little to stop it operating under time pressure in the room. The reliable move is structural — change the process so the impression has fewer openings, not the interviewer's mindset.
If you cannot say what specific behaviour earned a candidate their teamwork score, you are probably scoring your rapport with them. 'I liked them' is a similarity signal, not an assessment. Structure exists to make the difference visible.
What structure narrows — and how
Structured behavioural assessment is not one technique but three disciplines applied together. Each closes a specific channel that bias uses. None of them makes hiring mechanical; they concentrate human judgement on weighing evidence and take it away from resisting familiarity in real time, which is the thing human judgement is worst at.
The same probes for everyone
Ask every candidate the same behavioural questions, in the same order, chosen in advance from the competencies the role actually needs. This sounds obvious and is routinely skipped, because an unscripted interview feels more natural and more insightful. It is neither. When questions vary by candidate, you are no longer comparing like with like — you are comparing the answer a candidate happened to be asked with one you improvised because you were curious, or bored, or already sold. Standard probes make answers comparable, which is the precondition for scoring them fairly. Our structured interviews guide covers the mechanics of building and running a fixed question set.
Anchored scales, not adjectives
A behaviourally-anchored rating scale defines each score by observable behaviour rather than a vibe. 'Collaboration' stops being a gut read from one to five and becomes a rubric: this is what a strong answer looks like, this is adequate, this is weak. Anchors do the work that instinct cannot — they raise inter-rater agreement, meaning two fair reviewers converge on the same score, because they are applying the same definition instead of two private ones. The single most important discipline layered on top is scoring before discussion: independent scores captured before the panel talks stop the most senior or most confident voice from anchoring everyone else. A rubric read aloud after the room has already agreed is decoration, not structure.
Observed behaviour over claimed behaviour
The strongest structural move is to shift weight from what a candidate says they did to what you can watch them do. A behavioural interview is self-report; a work sample is behaviour. When a candidate handles a realistic task, gives feedback on a flawed piece of work, or navigates a scenario built from a real incident, you are rating observed conduct rather than a rehearsed narrative. This is where the AI Sandbox fits: realistic, job-shaped work that produces a reviewable artefact instead of a story about one. Our work-sample tests guide makes the fuller case; the short version is that observed behaviour is far harder to fake than claimed behaviour, and it favours no particular storytelling style.

Put together, these three disciplines narrow the channel bias runs through. Similarity bias has less room when rapport is not the medium of assessment; the halo effect has less room when each behaviour is scored independently against its own anchor; self-serving narratives have less room when observed behaviour carries the weight. The gains here are real and consistent — anchored, structured behavioural assessment is well established as more predictive and more reliable than unstructured judgement, and the structure, not the interviewer's instinct, is what earns it.
What structure does not fix
A credible account of behavioural assessment has to name its own weaknesses, because this pillar has two the others do not, and pretending otherwise is how a fair-sounding process ends up unfair. The honest position is that structure narrows bias; it does not close the channel, and it introduces one risk of its own.
Self-report can be gamed
Behavioural questions ask candidates to describe their own past behaviour, and self-report is inherently rehearsable. The STAR format is now so widely coached that a polished 'time I resolved a conflict' story tells you as much about interview preparation as about the candidate. Some answers are borrowed — a team's achievement told in the first person. Some are shaped to what the candidate thinks you want to hear. A candidate who interviews well is not necessarily a candidate who behaves well; the two skills overlap only partly, and mistaking the first for the second is a bias toward the articulate and the well-coached, which correlates with background.
The mitigations are partial and worth using anyway. Probe for specifics the rehearsed version glosses — what exactly did you say, what did the other person do next, what would you change. Score the reasoning and the observed detail, not the polish of the delivery. And lean on observed behaviour wherever you can: a work sample or a scenario response is a behaviour you watched, not a claim you have to take on trust. None of this makes gaming impossible. It raises the cost of it and lowers the reward, which is the honest most you can promise.
Anchors can encode one working style
The subtler risk is in the anchors themselves. A behaviourally-anchored scale is only as fair as the behaviour it rewards, and it is easy to write anchors that encode one culture's idea of the right way to work. 'Speaks up assertively', 'takes charge of the room', 'pushes back directly' — these describe a working style, not a job outcome, and they systematically favour candidates whose cultural or personality norms match the writer's. An anchor that rewards a particular style of confidence will screen out people who are quietly effective, and it will do it consistently enough to look like a real signal.
The mitigation is to anchor on outcomes, not styles. 'Resolved the disagreement and preserved the working relationship' is job-relevant and style-neutral; 'was assertive in the disagreement' is a preference dressed as a criterion. Draft anchors from real incidents in the role rather than from an idealised interviewer, review them across a diverse panel so a single author's norms do not become the standard, and rotate reviewers so no one perspective sets the bar. Even then, you verify by measurement rather than assumption — which is the point of the next section.
A perfectly consistent rubric can be consistently biased. If every reviewer applies the same anchor and that anchor rewards one working style, structure will produce uniform unfairness — and it will look rigorous while doing it. Consistency is necessary, not sufficient.
Verify the result, don't assume it
Structure is the design half of reducing bias; monitoring is the audit half, and behavioural assessment needs both precisely because its own biases can hide inside a consistent process. Adding a rubric does not guarantee fair outcomes — a stage can be perfectly standardised and still skew, whether from anchors that encode a style or from something upstream. Only outcome data tells you which. Track selection rates by group at the behavioural stage, apply the four-fifths rule, and treat a flag as a reason to inspect your anchors and probes, not as a verdict.
This is where the whole claim comes together. Structured behavioural assessment reduces bias relative to unstructured interviews because it gives similarity bias and the halo effect fewer openings; adverse-impact monitoring confirms the reduction actually happened rather than assuming it. Our explainer on adverse impact and the four-fifths rule covers the measurement in depth, and it is part of the answer here, not an optional extra. Fairness you cannot evidence is fairness you cannot defend — to a candidate, an auditor, or yourself.
When a behavioural stage flags under the four-fifths rule, read your anchors first. A style-encoding anchor is the most common hidden cause, and it is invisible until the outcome data points at it — because every reviewer was applying it faithfully.
Behavioural, in the context of the other pillars
Behavioural assessment is one of the five pillars of a hire, and it is strongest as one signal among several rather than a single gate. Its particular weakness — reliance on self-report — is exactly the weakness that demonstrated-skill pillars do not share, which is the argument for skills-based hiring as the foundation and behaviour as a complement. Situational judgement sits close to it; where behavioural interviews ask what you did, situational assessment asks what you would do. The pattern across all of them is the same: same evidence for every candidate, same rubric, monitored outcomes.
H-Evaluate is an AI-native skills assessment platform built to make that pattern the path of least resistance. Behavioural and situational assessments are generated per job description against a consistent rubric, so every candidate meets the same job-related probes rather than a recruiter's improvised questions. AI Sandbox work samples surface observed behaviour rather than claimed behaviour, and selection data is captured per role and per stage, so adverse-impact monitoring is a by-product of how you hire rather than a project you run under deadline. Structure narrows the channel; measurement proves it did — and the two together are how behavioural assessment earns its place.
The point of structure is not to make behavioural assessment scientific. It is to make 'I liked them' show up as what it is — so that liking someone is a thing you notice, not a thing you score.
Written by
Aayesha Patel · Co-founder, Hanzomon Inc
Co-founder of Hanzomon. Writes about skills-based hiring, fair assessment and building a better candidate experience.