AI-generated skills test
Prompt engineering test
Prompt engineering is the AI-hiring role most likely to be filled on faith. The title attracts certificate collectors and template resellers whose portfolios are tidy prompts that worked once, for something, somewhere — none of which tells you whether their approach holds up across a hundred real cases or survives the next model version. A CV cannot separate a fluent vocabulary from a real skill. A Prompt engineering test can, because it stops asking candidates to define "few-shot" and "chain-of-thought" and instead watches them work a failing prompt until they have measurably improved it.
H-Evaluate's Prompt engineering test measures the discipline underneath the jargon: systematic iteration, building evaluation sets, and writing instructions robust enough to survive a model change. Candidates read a set of failures, form a hypothesis about why the output breaks, change one variable, and check the fix held across the whole case set rather than the one example they were staring at. Every candidate works the same task under the same conditions, so process beats polish and a self-taught practitioner can out-prove a course badge. It maps to the AI Fluency pillar of H-Evaluate's five-pillar framework, run in the AI Sandbox with the model genuinely available and the diagnostic loop observed, not just the final artefact. Items are AI-generated for the specific role rather than reused from a bank.
What it measures
Systematic iteration
Reading a failure, forming a specific theory about why it happens, changing one variable and measuring the result — rather than tweaking wording at random until something passes. This measurement loop is the single strongest signal that separates a real practitioner from a template reseller.
Building and using evaluation sets
Turning "it feels better" into a case set with expected outputs, so improvement is a number rather than a vibe. Strong candidates build or ask for an evaluation set almost immediately and refuse to judge a change on one example.
Instruction robustness
Writing prompts and system messages that survive a model version change instead of over-fitting to today's model's quirks. Signals someone who builds for durability over cleverness and knows a prompt tuned to one release will break on the next.
Reading failure modes
Distinguishing hallucination from format drift from an over-eager refusal quickly, because the fix for each is different — and knowing when a problem is a retrieval or data issue the instruction layer cannot solve at all.
Question formats
Who it's for
Use this test for roles where prompt quality is close to the product: dedicated prompt engineers, AI and AI-agent engineers who own the instruction layer, and product or content staff running AI systems at scale. It is most useful once your prompts number in the hundreds and drift measurably when a model updates; below that, screen for broad AI fluency instead. Pair it with our AI fluency test for judgement and verification across tasks, and with a coding or domain test where the role also builds the surrounding system. It suits mid-to-senior hiring, since the value is in judgement rather than juniority.
How to read the results
- 1Score the loop, not the artefact. A candidate who reached a good prompt by luck should band below one who reached a slightly worse prompt through a clean, repeatable process — because next quarter, on a problem you cannot foresee, it is the loop that ships.
- 2Weight measurement discipline highest: did they build or ask for an evaluation set, change one thing at a time, and check the fix across the set? Random rewording that happened to pass is a weaker signal than a disciplined process that landed slightly short.
- 3Calibrate the bar to scope and seniority. Owning the prompts for one feature is a different demand from owning the agent-instruction strategy for a product line; read the band against what the role actually requires.
- 4Treat the result as one signal among several and pair it with a structured interview about a past evaluation set, a model update that broke their prompts, and a case they judged was not a prompt problem at all — process claims are easy to make and worth probing.
AI-generated skills test
Evaluate candidates on this skill with AI-generated questions
Configure a role-tuned assessment and watch it adapt by seniority — no signup.
Related roles
Related reading
How to hire a prompt engineer: test, don't take on faith
How to hire a prompt engineer in 2026: whether you still need the role, how to spot a template reseller, and the work sample that reveals real skill.
Prompt engineering by role: what good looks like
Prompt engineering by role is a core hiring skill, but 'good' differs for an engineer, analyst, support agent, SDR, PM or recruiter. A per-vertical hub guide.
Frequently asked questions
What does a prompt engineering test measure?
It measures the discipline behind good prompting rather than clever phrasing: systematic iteration, building evaluation sets, and writing instructions robust enough to survive a model change. The core question is whether a candidate can turn "the output feels off" into a measured, repeatable improvement — reading failures, changing one variable, and checking the fix held across a whole case set. Vocabulary like "few-shot" is the smallest part of the job.
How do you test a prompt engineer?
Give them a mediocre prompt and a set of cases it fails on, then watch them diagnose and iterate — the job compressed into an exercise. You are grading the loop, not the final prompt: do they read the failures first, group them by type, change one thing and re-run, and build or ask for a way to measure the change across the whole set? Strong candidates reach for an evaluation set almost immediately; weak ones reword at random or paste a template.
Are prompt engineering tests reliable for hiring?
A well-built one is far more reliable than a CV or a definitions quiz, because it observes the actual behaviour the job needs rather than a candidate's account of it. Reliability comes from scoring an observable process — iteration, measurement, robustness — against a consistent rubric for every candidate, and from running it with the model genuinely available so you see the loop, not just the artefact. Pair it with a structured interview to confirm the process generalises.
Do you still need to test prompt engineering in 2026?
Where prompt quality is genuinely the product — support automation, content systems, agent instructions at scale — yes. Writing a passable prompt is now table stakes for many jobs, but owning an instruction layer of hundreds of prompts that drift when a model updates is a specialism worth testing for directly. If your AI surface is a feature or two, screen for broad AI fluency instead of a dedicated prompt-engineering skill.
What is the difference between a prompt engineering and an AI fluency test?
An AI fluency test measures the general judgement of working well with AI across any task — delegating sensibly, verifying output and using it responsibly. A prompt engineering test goes deeper into the instruction layer specifically: iterating systematically on prompts, building evaluation sets, and making instructions robust to model changes. Fluency asks whether someone works well with AI; the prompt test asks whether they can engineer and measure the instructions that shape it.