AI-generated skills test

AI fluency test

By 2026 most knowledge work happens with an AI assistant in the loop, which makes AI fluency one of the most predictive things you can screen for — and one most hiring processes still measure badly. A CV line about being "AI-savvy" tells you nothing, and asking a candidate how they would use AI responsibly earns you the textbook answer every time. Description-only screening fails here because fluency is a behaviour, not a claim. An AI fluency test replaces the account with the act: it puts a candidate in a live task with AI available and observes how they actually work with it.

H-Evaluate's AI fluency test measures judgement, not enthusiasm or prompt trivia. The strongest single signal is not how eagerly someone reaches for AI but how well they know its limits — so the test is built around a task where the model is confidently wrong, and you watch whether the candidate catches it. It reads capability across four dimensions rather than one generic score, and calibrates the bar to the role and seniority. It maps to the AI Fluency pillar of H-Evaluate's five-pillar framework, run in the AI Sandbox so the model is genuinely available and both the outcome and the process are captured — because two candidates can reach the same answer, one by spotting the error and one by luck. Tasks are AI-generated for the specific role rather than pulled from a shared bank.

What it measures

Delegation

Deciding what to hand to AI and what to keep by hand — using the model where it genuinely raises quality and choosing not to use it where doing the work directly is better. For senior roles this becomes the judgement of splitting a hard problem sensibly and justifying the split.

Description

Telling the model what is actually wanted — framing a task, giving the right context and constraints, and steering toward a useful result. Not clever phrasing for its own sake, but the clarity that makes AI produce work worth building on.

Discernment

Judging what comes back: catching a confidently-wrong answer, knowing which claims to check hardest, and verifying before relying on output. The single highest-signal dimension, surfaced by a task with a plausible-but-false result planted in it.

Diligence

Using AI responsibly and owning the result — handling confidential data and potential bias sensibly, and treating accuracy as the candidate's responsibility rather than the tool's. Watched through a scenario that invites, and rewards resisting, a careless shortcut.

Question formats

Hands-on task in the AI Sandbox with the model available and the process observedConfidently-wrong item — a plausible-but-false or subtly broken output the candidate must catchRealistic scenario questions about how a tool behaves and where its output needs checkingData-sensitive case testing responsible handling of confidential or biased materialDelegation judgement — deciding which parts of a task to hand to AI and which to keep by handShort written response explaining why a change or check was made, to read the reasoning behind the action

Who it's for

Use this test for any role now done with AI in the loop — which is most of them: engineers, analysts, support, sales, marketing and operations staff who use these tools daily. It is not an engineers-only measure. Calibrate the bar to the role: safe defaults and diligent checking for a junior hire, good delegation decisions under ambiguity for a senior one. Weight the pillar down where a role genuinely uses AI little, rather than inventing a task. Pair it with our prompt engineering test where the role owns an instruction layer, and with domain and situational tests to build the full picture across the five pillars.

How to read the results

  • 1Read it as role-relative, not as one universal AI-fluency score. The bar for a junior support agent drafting a reply with AI help is not the bar for a senior engineer weighing whether to trust a generated migration; band the result against what the specific role demands.
  • 2Score behaviour, not the impression of savviness. "They accepted a confidently-wrong output without checking it, on the same task every candidate received" is defensible; "they seemed comfortable with AI" is not — confidence and jargon are exactly what you do not want to reward.
  • 3Weight the four dimensions to the job, and read them separately: a candidate strong on Description but weak on Discernment ships fluent, unverified errors, which is a targeted area to probe in the interview.
  • 4Treat it as one signal among several. It sits as one weighted pillar, not a bolt-on badge — pair it with a structured interview and the other pillars, and give every candidate the same tool access so you measure fluency, not who had a paid subscription at home.

AI-generated skills test

Evaluate candidates on this skill with AI-generated questions

Configure a role-tuned assessment and watch it adapt by seniority — no signup.

Related roles

Related reading

Frequently asked questions

What does an AI fluency test measure?

It measures the judgement of working well with AI, not enthusiasm or prompt trivia — across four capabilities: delegating the right work to AI, describing tasks clearly, discerning whether output can be trusted, and using it responsibly. The strongest signal is how well a candidate knows the tool's limits: whether they catch a confidently-wrong answer and verify before relying on it, rather than how eagerly they reach for AI in the first place.

How do you assess AI fluency in hiring?

Give candidates a realistic task with AI available and observe how they use it, rather than asking them to describe it — description-only assessment earns the textbook answer from almost everyone. The highest-signal item plants a plausible-but-false or subtly broken output and watches whether they catch it. Capture the process, not just the final artefact, and calibrate the bar to the role and seniority instead of chasing one generic score.

Should candidates be allowed to use AI during the test?

Yes, if the job uses AI. Banning it measures a version of the role that no longer exists, so it cannot predict on-the-job behaviour. Give candidates access to AI tools and observe where they rely on the model, where they verify, and where they choose not to use it — that observation is the assessment. H-Evaluate runs the task in the AI Sandbox so the model is genuinely available and the process is captured.

Is AI fluency the same for every role?

No — it is role-relative. The bar for a junior support agent, who drafts a reply with AI help and checks it before sending, is not the bar for a senior engineer deciding whether to trust a generated migration. Assess fluency against what the specific role demands, calibrated to seniority, which is why it sits as one weighted pillar rather than a single universal score, and why the same task can be pitched at two difficulty levels.

What is the difference between an AI fluency and a prompt engineering test?

An AI fluency test measures general judgement working with AI across any task — delegating, describing, discerning and using it responsibly. A prompt engineering test goes deeper into the instruction layer: iterating systematically on prompts, building evaluation sets, and making instructions robust to model changes. Fluency asks whether someone works well with AI in their role; the prompt test asks whether they can engineer and measure the instructions that drive an AI system.