Hiring · August 2, 2026 · 10 min read
Prompt engineer job description template (free, 2026)
A prompt engineer job description template for 2026, plus the AI fluency section every other template skips and the requirements you can actually assess.
← Part of The five pillars of hiring: what assessments measure
On this page
This page is for the hiring manager, founder or talent partner who has decided to write a prompt engineer job description and wants one that survives contact with real candidates. The role owns the instruction layer between your product and the model — the prompts, system messages and agent instructions — plus the evaluation sets that prove those instructions work. Writing the description is hard for three reasons. The title is young enough that no two companies mean the same thing by it. The market boilerplate is thin, dated, and copies itself. And the applications will arrive stuffed with prompt collections and course certificates that look impressive and predict almost nothing. This role sits wherever prompt quality is genuinely the product — support automation, content systems, agent instructions running at scale — usually inside engineering or applied AI. Below is a copy-paste template, plus the two sections every competitor leaves out: what to assess, and what AI fluency actually means here.
Before you post this: is prompt engineering still a standalone job in 2026? Our position is that it has consolidated, not vanished. It is now a systems role — evaluation, retrieval design, model-behaviour debugging — not a wordsmithing one. If what you need is clever phrasing, you do not need this hire. If you need measured output quality at scale, read on. The how to hire a prompt engineer guide covers the deciding-whether question in full.
The prompt engineer job description template
Copy the sections below and replace the bracketed placeholders. Keep the AI fluency section — it is the part that separates a description candidates respect from one they can game.
About the role
[Company] is hiring a prompt engineer to own the instruction layer of [product or system] — the prompts, system messages and agent instructions that shape how our models behave in production. You will treat prompts as code: versioned, tested against evaluation sets, and measured rather than eyeballed. This role sits within [team] and works closely with [engineering / product / operations] to turn 'the output feels off' into measured, repeatable improvements.
Responsibilities
- Own the prompts, system messages and agent instructions for [product surface], and the evaluation sets that keep them honest.
- Diagnose output failures systematically: reproduce the problem, group failures by type, form a hypothesis, change one variable, and measure the result against a case set.
- Build and maintain evaluation sets with expected outputs, so every change is a number rather than an impression.
- Version prompts like code — with change history, rollback, and a record of which change moved which metric.
- Catch and diagnose model-behaviour regressions when a model version or family changes, and put guards in place so the next one is caught automatically.
- Decide when a problem lives in the instruction layer versus retrieval, data, or the model itself, and route it accordingly.
- Partner with engineering on retrieval and tool-call design where those decisions shape model behaviour.
- Report output quality against defined metrics to [stakeholders], in plain language a non-technical owner can act on.
Requirements
- Demonstrated systematic iteration: you can show a case where you turned a vague quality complaint into a measured, repeatable improvement.
- Experience building evaluation sets and reasoning about output quality as a number, not a vibe.
- Fluency with common failure modes — hallucination, format drift, over-eager refusal — and the different fix each one needs.
- Judgement about the limits of the instruction layer: you say 'this is a retrieval problem, not a prompt problem' when it is true.
- Comfort working across at least two model families and an understanding of how behaviour differs between them.
- Clear written communication — you can explain a trade-off to a non-technical owner without hiding behind jargon.
- Enough engineering literacy to work inside a versioned codebase and reason about the system around the model.
Nice to have
- Experience owning prompts for a live product at scale, where drift carried a real cost.
- Background in applied machine learning, data analysis, or a numerate discipline.
- Familiarity with retrieval-augmented systems and agent frameworks.
- Experience building or maintaining an evaluation harness that ran unattended.
AI fluency expectations
This is the section no ranking template in the field includes, and for a prompt engineer it is the deepest in this series — the job is AI fluency. Paste these verbatim:
- Evaluates model output systematically, against a case set with expected results, rather than declaring a prompt 'better' because it reads better.
- Versions prompts like code — with history, rollback, and a clear record of which change moved which metric.
- Knows the failure modes of at least two model families and how the same prompt behaves differently across them.
- Detects subtle model-behaviour regressions after a version change, and verifies a fix held across the whole set before declaring victory.
- Uses AI tools to move fast while checking their output against the real system — delegating what is safe to delegate, and catching a confident answer that is quietly wrong.
We checked: not one ranking prompt engineer job description template in the field includes an AI fluency expectations section. For the role whose entire job is AI fluency, that is a strange omission — and it is exactly the section that filters a systematic practitioner from a template reseller. Keep it. It is the part of this template competitors cannot copy without understanding the role.
What we offer
[Compensation range], [equity if applicable], and [benefits]. You will own a lever that visibly moves [the metric that matters], with the evaluation infrastructure and autonomy to move it. [Working arrangement — remote, hybrid, location]. [One honest line about the team and how you work.]
How do you adapt this template?
Turn the seniority dial with scope, not with a longer requirements list. Someone owning the prompts for a single feature is a different hire from someone owning the agent-instruction strategy for a product line — scope and pay them accordingly. Cut hard for a startup; expand the governance and reporting lines for an enterprise. Then delete the three things people wrongly copy from legacy job descriptions.
- A degree requirement. The strongest prompt engineers are often self-taught practitioners and career-switchers from numerate fields; a degree gate filters them out and predicts nothing.
- 'X years of prompt engineering.' The discipline is a few years old at most. Anyone claiming a decade is rounding up, and the requirement quietly rewards early adopters over better practitioners.
- A named certificate or course. Certificates measure attendance, not the measurement discipline this role runs on. Ask to see a measured improvement instead.
- A fixed list of specific tools or model versions. Model families change fast; hire for transferable judgement across them, not fluency in this quarter's stack.
For a startup, collapse the responsibilities to the core loop — own the instruction layer, build the evaluation sets, catch the regressions — and let one person carry it alongside adjacent work. For an enterprise, add the reporting, review and documentation lines the template hints at, and be explicit about which model families and surfaces are in scope. The how to write a job description guide covers the general mechanics of writing around observable behaviour.
What should you assess instead of trusting the CV?
Map every requirement bullet to something you can observe, because for this role the CV is close to noise. A prompt engineer's CV is a tidy library of prompts that worked once, for something, somewhere — and a template that produced a lovely demo output tells you nothing about whether it survives a hundred real cases or the next model version. Assess the loop, not the lexicon. The requirements above land on the five public pillars of candidate evaluation like this — capability level only, no scores:
- Cognitive: the diagnostic loop itself — reading failures, isolating a variable, reasoning about cause. This is the strongest single signal for the role.
- Domain: evaluation design, retrieval literacy, and knowing the limits of the instruction layer.
- Situational judgement: deciding whether a problem is a prompt problem at all, and what to escalate versus fix.
- Behavioural: measurement discipline over the temptation to declare victory on a lucky output.
- AI fluency: prompt versioning, cross-family failure-mode literacy, and verifying assistant-drafted work against the real system.
The way to assess all five at once is a job-shaped work sample: hand the candidate a mediocre prompt and a set of cases it fails on, then watch them diagnose and iterate. That is the job compressed into an hour, and it is a work sample test rather than a quiz. Because the role is intrinsically AI-native, observe how they work with the model directly — the 4D framework of AI fluency (Delegation, Description, Discernment, Diligence) gives you a rubric, and how to assess AI fluency walks through reading strong versus weak as you watch. Running the exercise in a realistic AI Sandbox, where the model is genuinely available and you observe the process rather than only the artefact, is where our AI-native skills assessment platform view lands. To calibrate what the instruction layer looks like across other jobs, prompt engineering by role and prompt engineering for software engineers are useful companions; the deciding-whether-to-hire logic lives in how to hire a prompt engineer, and if what you actually need is agent architecture, the AI agent engineer job description is the neighbouring role.

Illustrative weights — configurable per role, locked at the first candidate for comparability.
How do you spot a prompt engineer who cannot actually do the job?
This is the worst impostor problem in the whole AI-hiring landscape, and the honest position is worth stating plainly. Because the role once had a hype cycle, it attracts two kinds of pretender. The certificate collector, who has completed several 'prompt engineering' courses and can recite frameworks with acronyms. And the template reseller, whose entire portfolio is a copy-paste library. Both interview fluently. Neither can turn 'the output feels off' into a measured improvement. We sell candidate evaluation for a living, so weigh my framing accordingly — but the tell holds regardless of who says it: ask how they know a prompt got better, and listen for a number against a case set rather than 'it reads better.' The red flags in an application:
- A CV that lists prompt collections and frameworks but no measured output improvement anywhere.
- Portfolio prompts with no evaluation set behind them — outputs that looked good once, with no evidence they hold across cases or model versions.
- Fluent vocabulary ('few-shot', 'chain-of-thought') and no story about a model update breaking their prompts and what they did next.
- Certificates standing in for demonstrated judgement, and a discomfort when asked to work a live failing-case set.
- Claims that a prompt can fix everything, with no instinct for when a problem is really retrieval or data.
You may not need this role at all. If your AI surface is a feature or two, your existing engineers or a product manager with real AI fluency can own the prompts alongside their other work. And do not hire a prompt engineer to paper over a model or data problem — if the outputs are bad because retrieval is broken, no prompt will save you, and the specialist will spend six months politely explaining that. Diagnose where the fix lives before you post the role.
The threshold is not 'do we use AI' — everyone does now. It is 'does one person moving output quality a few points change a business metric we care about.' Support automation where a small drop in resolution quality shows up in churn; a content system generating at volume where drift is a brand risk; agent instructions running unattended where a bad edge case costs money. If prompt quality is genuinely the product, hire the specialist and assess for the loop. If it is a shared skill, hire for AI fluency across the team instead — the skills-based hiring guide and a structured interview loop will serve you better than a specialist title you do not need.
A prompt engineer job description that reads like a spec you can test against will out-hire ten polished templates that read like a wish-list. Write it around measured outcomes, keep the AI fluency section every competitor skips, and assess the loop — not the lexicon — or you will keep mistaking a good vocabulary for a real skill.
Written by
Aayesha Patel · Co-founder, Hanzomon Inc
Co-founder of Hanzomon. Writes about skills-based hiring, fair assessment and building a better candidate experience.