All posts

Technology · July 21, 2026 · 10 min read

AI Sandbox assessment: how to run one and score it

Your hires will use AI from day one, so assess them with it. What a good AI Sandbox task looks like, what to measure, and how to read the signals fairly.

By Jakir Patel · Founder, Hanzomon

Share

Part of AI-generated assessments: the complete 2026 guide

Technology
On this page

AI-native hiring starts from an honest premise: your candidates have AI, and so will your hires from day one. This guide is for hiring managers and assessment designers who have accepted that and now need something practical — a task that produces real signal rather than a policy that produces a false sense of control. The AI Sandbox is how you turn candidate AI use from a threat into the thing you are actually measuring. If the concept is new to you, the short explainer covers what it is; this is the deep-dive on running one well and reading the result without fooling yourself.

The stakes are worth stating plainly. Every assessment format that assumes an isolated candidate is now measuring a situation that will never recur after the offer letter. That does not just make the score less useful — it makes it misleading, because it is precise about a skill the role no longer isolates. The question that predicts performance is no longer whether someone can solve a problem unaided; it is whether they can solve it well with a capable, confidently wrong assistant in the loop.

What does a sandbox task look like?

A sandbox task is open-ended, realistic, and set in a live environment with the tools of the job available — including an AI assistant. The candidate is not answering questions about the work; they are doing it. A support example: here is a prompt that summarises meeting notes but keeps leaking who-said-what into a public channel — fix it so the summary is structured and free of personal attribution. An engineering example: this function has a subtle bug and you have an AI assistant — get it right and be ready to explain what you changed.

Notice what both examples have in common. There is a recognisable good outcome, so the task is gradable. There is a trap that rewards reading carefully rather than prompting quickly. And there is nothing to look up, because the task is about this artefact rather than about general knowledge. That last property is what makes a sandbox resistant to the leak problem that eventually catches every fixed question bank.

The AI Sandbox in action — a live task where the candidate works with AI tools while their process is assessed.

What are you actually measuring?

The finished artefact matters, but it is not the whole grade — two candidates can hand in similar output and be very different hires. What you are really scoring is the collaboration, across a few observable behaviours:

  • Prompting and direction — do they steer the tool clearly toward a real, specific outcome, with the right constraints and exclusions, or fire off vague requests and hope?
  • Verification instincts — do they check the output, or ship it unread? Catching a wrong answer, a subtle bug, or a confidently-stated fact that is not true.
  • Judgement and correction — when the first result is flawed, do they refine the prompt, fix it by hand, or recognise that the AI is the wrong tool for this part?
  • Quality of the finished work — does the end result actually clear the bar for the role, produced in the realistic, tool-assisted way the job really runs?

The shift is this: you are grading the collaboration, not just the artefact. The candidate who reached a good result by verifying and correcting is a safer hire than one who reached a similar result by luck and never looked twice.

These four map onto the same capabilities the 4D framework uses to describe AI fluency generally — delegation, description, discernment and diligence. The sandbox is where those stop being a vocabulary and become something you can watch happen.

Strong and weak signals

Across real sandbox tasks, the difference between a strong and a weak candidate is usually visible in how they work, not just in what they hand in. The strong signals:

  • Diagnoses why the first attempt failed before trying again, rather than resubmitting near-identical prompts.
  • Writes clear instructions with explicit inclusion and exclusion rules, and a stated output format.
  • Catches the AI's mistakes — a wrong fact, a missed edge case — and fixes them.
  • Ships work that would hold up in front of a customer or a code reviewer.
  • Can narrate what the tool produced and why they accepted or rejected it.

The red flags:

  • Pastes AI output unchanged and moves on, with no evidence of having read it.
  • Vague, single-line prompts with no constraints, and no reaction when the result is wrong.
  • Cannot explain what the AI produced or why they accepted it.
  • Confident, polished output that quietly contains an error they never noticed.

The last one is the expensive failure, and it is the reason output-only grading does not work. Unverified AI output is very good at looking finished. A rubric that rewards polish will systematically prefer the candidate who did not check their work, which is the precise opposite of what you were trying to buy.

Where sandbox assessments go wrong

The format is not self-justifying, and there are three ways teams get a sandbox that produces confident noise. The first is a task with no failure mode: if the obvious approach works, every candidate looks careful and the rubric has nothing to separate. The second is grading from a transcript alone. Reading a log of prompts tells you what someone typed, not whether they understood what came back, and it quietly rewards people who narrate their thinking over people who simply do the work well. The third and most common is scope creep — the task grows until it is a take-home in disguise, at which point you have reintroduced every problem you were trying to escape.

There is also a subtler trap around tool familiarity. If the task can only be completed efficiently with one specific assistant, you are measuring interface knowledge rather than judgement, and you will systematically prefer candidates who happen to use the same tools you do. The fix is to keep the assessable behaviours tool-agnostic: whether someone checked the output is a question you can answer regardless of which model produced it.

If a candidate could pass your sandbox task by pasting the prompt into any assistant and submitting the first response unread, the task is not measuring what you think it is. Test it against that behaviour before you send it to anyone.

Why let candidates use AI at all?

Because banning it tests a scenario your hires will never face. The old instinct — lock the environment down, add proctoring — measures whether someone can work in a way they never will again. Meanwhile the real risk of the AI era is not that people use these tools; it is that they ship the tools' mistakes without noticing. Verification is the core skill now, and the only way to measure it is to let the AI into the room. It is the same logic behind why proctoring cannot stop cheating with AI — you do not beat candidate AI by banning it, you out-design it by assessing it.

Freshness — nothing to look up
Behavioural flags
AI-answer detection
Proctoring (optional, consented)

Layered defence: freshness removes the payoff, and each signal narrows what slips through.

There is a secondary benefit that rarely gets mentioned: candidates prefer it. Strong senior people increasingly decline assessments that feel hostile or artificial, and a task that hands them real tools and asks them to do real work reads as respect rather than suspicion. The format that produces better signal also produces fewer declines, which is not a trade-off you often get.

None of this means integrity stops mattering. It means the integrity question changes shape: instead of asking whether a candidate used AI, you ask whether the person who did the work is the person you are hiring, and whether the evidence in front of you is theirs. Those are answerable with far less intrusion than a locked-down proctoring stack, and they hold up better, because they do not depend on preventing something that is now ambient.

How do you design a good sandbox task?

  • Make it job-relevant — a real slice of the work, not a puzzle chosen because it is easy to grade.
  • Give real tools, including AI, so the workflow matches the job.
  • Keep it open-ended but with a clear bar — there should be a recognisable good outcome.
  • Plant something worth catching, so verification behaviour has a chance to show itself.
  • Time-box it humanely — enough to show process, not a marathon that filters for free time.
  • Score every candidate on the same rubric so judgement-heavy answers stay comparable.

The fourth point is the one teams most often miss. If nothing in the task can go wrong, verification is invisible and every candidate looks equally careful. A task with a plantable flaw — a subtly incorrect assumption, an edge case the obvious approach misses — is what separates the person who reads output from the person who forwards it. This is the sandbox's version of what makes any work sample predictive in the first place: realism with a consequence attached.

Fairness and how to read the result

Because a sandbox is open-ended, consistency is what keeps it fair: the same rubric for everyone, and a human in the loop to interpret rather than a black box deciding alone. That is not only good practice — for automated employment decisions it is increasingly a legal requirement, and an open-ended format needs a documented rubric precisely because it has more room for inconsistency than a multiple-choice test does.

Resist over-indexing on output polish. A beautiful artefact produced by unquestioned AI is worth less than a rougher one produced with real judgement, and a rubric that cannot express that difference will keep making the same expensive mistake. Used well, the AI Sandbox pairs naturally with AI fluency as a scored pillar rather than standing alone: the sandbox gives you the behavioural evidence, the pillar gives you something comparable across candidates.

One practical habit makes the whole format more defensible: write down what a strong, adequate and weak response looks like before the first candidate sees the task. Rubrics written after the fact drift toward whatever the first impressive candidate happened to do, and that is how an open-ended assessment quietly becomes a preference for people who work the way the reviewer does. Deciding the bar in advance also forces you to check that the bar is reachable, which is a useful thing to discover before you have rejected three people against it.

Finally, calibrate reviewers against each other. Have two people score the same two or three sessions independently and compare, before either of them scores a real candidate alone. Where they disagree, the disagreement is almost always about what counts as sufficient verification — which is exactly the conversation you want to have once, deliberately, rather than repeatedly and invisibly across a hiring round. This is the same discipline that makes structured interviews outperform unstructured ones, applied to a format that needs it more, not less.

Use the sandbox output as the agenda for the interview that follows. Walking a candidate through their own session — why that prompt, what made you check that, what would you verify before shipping — turns a subjective format into a documented, defensible decision.

If you want to see one built for a role you are actually hiring, you can request a sample assessment, or book a demo and watch a role-tuned assessment composed from a job description you bring.

Stop asking whether a candidate can work without AI. Hand them the tools they will actually use, and watch how well they work with them.
AI SandboxAI fluencyAssessment designHiring integrity
J

Written by

Jakir Patel · Founder, Hanzomon

Building H-Evaluate — AI-native, quality-gated hiring assessments. Writes about assessment engineering, hiring integrity and compliance-first AI.

Put this into practice

The assessments, role guides and calculators that turn what you have just read into a hiring decision.

Frequently asked questions

What is an AI Sandbox assessment?

A live, hands-on task where a candidate works on a realistic problem with AI tools available, and is evaluated on the collaboration — how they prompt, verify and correct, and whether the finished work is actually good — rather than on whether they can work without AI. It measures working with AI instead of testing in an artificial, AI-free bubble.

How do you grade an AI Sandbox task fairly?

Score the process against a consistent rubric — verification behaviour, prompt craft and domain grounding — not just the final artefact. Apply the same rubric to every candidate and pair it with human review, so a polished-looking output does not outweigh sound judgement. Consistency is what keeps an open-ended task defensible when someone challenges the result.

Will candidates not just let the AI do everything?

That is exactly what the task surfaces. Someone who pastes AI output unchecked scores poorly on verification and correction; someone who directs the tool, catches its mistakes and ships something they can stand behind scores well. The AI doing the typing is the point — the judgement around it is what you are measuring.

How long should an AI Sandbox task take?

Long enough to show process, short enough that it does not filter for free time. Most roles are well served by a task in the thirty-to-sixty-minute range, because the behaviours you care about — how someone reacts when the first result is wrong — appear early. A multi-hour task mostly measures who could afford to spend the afternoon.

Does an AI Sandbox replace a technical interview?

No, it replaces the part of the interview that was a poor proxy for the work. The sandbox gives you evidence of how someone actually operates; the interview is where you walk through that evidence with them and ask why they made the calls they made. Used together, the interview stops being a memory test and becomes a conversation about real work.

Is it fair to candidates who have not used AI tools much?

It is fair in the same sense that testing spreadsheet skills is fair for an analyst role: it measures something the job genuinely requires. What keeps it fair in practice is assessing judgement rather than tool trivia, so someone who reasons carefully about output they did not write is not penalised for being unfamiliar with one particular assistant's interface.

Related posts

See it on your own job description

Join the early-access waitlist and watch H-Evaluate build an assessment for a real role.

See it on your own job description