All posts

Hiring · August 2, 2026 · 10 min read

AI Product Manager Job Description Template (2026)

A free AI product manager job description template for 2026, plus the AI-fluency and evaluation sections most templates leave out. Copy, adapt, and hire.

By Aayesha Patel · Co-founder, Hanzomon Inc

Share

Part of The five pillars of hiring: what assessments measure

Hiring
On this page

If a team at your company now ships a feature that talks back — a copilot, a summariser, an agent that takes actions — you are probably writing an AI product manager job description, and it is harder to write than it looks. This page is for the hiring managers, heads of product and recruiters who have to describe the role clearly enough to attract the right people and screen out the many who simply added 'AI' to a product title. An AI product manager owns features whose behaviour is probabilistic: they decide what a model should and shouldn't attempt, define what good output means when no two runs are identical, and own what happens when the model is confidently wrong. The title is new, the market boilerplate is thin and mostly dated to 2024, and every second CV now claims to have shipped an AI product. So the job description has to do real work — it has to read like something you can test a candidate against. Below is a copy-paste template, plus the two sections almost nobody else's template has: what to assess instead of trusting the CV, and the AI fluency expectations that separate the shippers from the rebranders.

The AI product manager job description template

Copy the blocks below and swap the [bracketed placeholders] for your specifics. The wording is deliberately concrete and current to 2026 practice — verb-first, testable, no filler. Cut what doesn't apply, but resist the urge to soften the responsibilities into generic product-manager copy; the whole point is that this role is different.

About the role

[Company] is hiring an AI product manager to own [product or feature area], where the core behaviour is driven by a model and will not behave the same way twice. You will decide what the model should and shouldn't attempt, define what good output means in checkable terms, and own the experience for when it gets things wrong. You will work day to day with [engineering / applied ML / design / data], and report to [manager]. This is [seniority level]; [on-site / hybrid / remote] from [location].

Responsibilities

  • Scope what the model should and shouldn't attempt, drawing the line between cases where a generated answer is genuinely useful and cases too risky or unreliable to ship.
  • Define evaluation criteria for each feature: what good output means in concrete, measurable terms, and how the team catches a regression before users do.
  • Set and hold a quality bar for non-deterministic output — decide how good is good enough to ship, knowing no version will be perfect.
  • Make build-versus-buy calls on models, weighing cost, latency, control and how fast the frontier is moving.
  • Own the guardrails and escalation experience: what the product does when the model is uncertain or wrong, and how a user reaches a human or a safe default.
  • Reason about error rates with non-technical stakeholders and translate model limits into an honest roadmap, not a cherry-picked demo.
  • Prototype and pressure-test feature ideas with AI tooling before committing engineering time.
  • Partner with [governance / legal / risk] on the responsible-use and compliance implications of shipping a probabilistic feature.

Requirements

  • A track record of owning at least one model-backed feature end to end — from scoping through launch and iteration — where the output was genuinely non-deterministic.
  • Demonstrated evaluation literacy: you can define what good output means for a feature and describe how you measured it in production.
  • Sound product craft — discovery, prioritisation, stakeholder alignment — applied to a component you cannot fully specify in advance.
  • Comfort reasoning about error rates and failure modes, and explaining them plainly to people who are not technical.
  • Judgement about when a model is the wrong tool, and the willingness to argue for a deterministic rule or a plain form instead.
  • Clear written communication — the role runs on evaluation docs, incident write-ups and honest roadmaps.

Nice to have

  • Experience owning the aftermath of a real model failure in production — a hallucination incident, a quality regression, a safety escalation.
  • Familiarity with [your domain], where the cost of a confidently wrong answer is well understood.
  • Hands-on comfort with evaluation tooling or building lightweight test harnesses for model output.

AI fluency expectations

This is the section legacy templates leave out entirely. It is not a tools checklist; it is a statement of the judgement the role demands. Paste these bullets and adapt the specifics to your product.

  • You can write evaluation criteria for model outputs — defining, for a given feature, what good looks like in terms concrete enough to measure and re-run.
  • You have enough prompt literacy to prototype and pressure-test a feature yourself, and to tell a genuinely robust behaviour from one that only worked in the demo.
  • You reason about error rates and failure modes out loud, and you design the product for the times the model is wrong rather than filing them as edge cases.
  • You can judge when a workflow should not use AI at all, and you make that call on the merits rather than reaching for a model reflexively.
  • You use AI tooling in your own work with discernment — delegating what it does well and verifying what it doesn't before it reaches a user.

The AI fluency expectations section is the part no ranking competitor template includes — we checked the field, and not one of them has it. It is also the part that does the most screening. A CV can claim an AI product; only concrete fluency bullets give you something to test the claim against.

What we offer

[Compensation range and equity], [benefits], and [the interesting bit: the product surface, the users, the model problems worth solving]. We assess candidates on demonstrated skill, not pedigree, and we tell you what the loop looks like before you start it. [Add your own culture and growth specifics — keep it honest and specific to this role.]

How do you adapt this template?

Dial the seniority up by raising the stakes of the probabilistic component, not the years-served bar: a junior owns a contained feature with a clear failure mode, a senior owns a model-central product where a confident mistake costs real money or trust. Cut ruthlessly for a startup; expand the governance and stakeholder lines for an enterprise. And delete the boilerplate that migrated in from legacy product-manager templates.

Startups should strip this to the responsibilities and the AI fluency expectations, then let one person wear several hats — the evaluation-and-guardrails core is non-negotiable, the rest can flex. Enterprises should keep the compliance, governance and stakeholder-communication lines and add their own review gates, because a model failure at scale is a public one. Either way, the responsibilities are the spine; keep those tight.

The lines people wrongly copy from old templates are the ones to watch. Three of them do active harm here:

  • A degree requirement. It filters out capable people and predicts nothing about whether someone can hold a quality bar on non-deterministic output. Drop it — the case for skills-based hiring is strongest exactly where the credential map is newest.
  • Fixed years-with-a-named-tool bars ('5+ years with [model or framework]'). The tooling turns over faster than any tenure clock; a rigid number screens for longevity, not judgement.
  • A wall of trendy model names. Listing this quarter's frontier models dates the post within weeks and rewards buzzword-matching over the evaluation skill you actually need.

Write the responsibilities as behaviours you could observe in a work sample, not aspirations. 'Define evaluation criteria for a feature' is testable in an afternoon. 'Passionate about AI' is not. If a bullet can't be assessed, it is decoration — cut it or rewrite it until it can.

What should you assess instead of trusting the CV?

A CV tells you what someone was in the room for, not what they can do — and in this role the titles are especially noisy. So map each requirement bullet to something you can actually observe. We think about candidate evaluation across five capability pillars: cognitive, domain, situational judgement, behavioural, and AI fluency. The template's requirements ladder onto them cleanly.

Domain
25%
Behavioural
20%
Situational
20%
Cognitive
15%
AI Fluency
10%
AI Sandbox
10%

Illustrative weights — configurable per role, locked at the first candidate for comparability.

  • Cognitive — the reasoning behind a build-versus-buy or model-versus-rules call, tested by asking them to make one and defend it.
  • Domain — real product craft plus the specifics of shipping on a model, tested by having them write evaluation criteria for a plausible feature.
  • Situational judgement — how they triage a live hallucination incident: immediate mitigation versus systemic fix, and who they think about first.
  • Behavioural — whether they own a past model failure plainly, including what it cost, or reach for blame.
  • AI fluency — how they use AI tooling in the work itself: what they delegate, and where they catch it being wrong before a user does.

The way to see all five is job-shaped work, not trivia. Ask candidates to write the evaluation criteria for a proposed feature, triage a hallucination incident, and reason through a model-versus-rules trade-off aloud — with AI tools genuinely available, because that is how the job is done. Our AI Sandbox is built to put a candidate in that kind of realistic environment and let you watch the process, not just read the artefact — which is, admittedly, exactly what a vendor would say. What you are grading is judgement under non-determinism, and you can only see judgement by watching someone exercise it. For the how, how to hire an AI product manager walks through the full loop and the exercises; how to assess AI fluency and the 4D framework cover reading the signals, with delegation and discernment carrying the most weight for this role, on the same show-me-don't-tell-me logic behind any real work sample.

A generated question set derived from a job description, showing how role requirements map to assessable tasks
A job description is a spec you should be able to test against. From the same role definition, a candidate evaluation can be generated per job — turning each requirement bullet into something observable rather than asserted.

Why does every AI product manager CV look qualified — and how do you tell them apart?

Because the title is new and the market rewards claiming it, almost every product-manager CV now says 'AI products.' Most of it is honest and most of it is thin. The gap you are hunting for is between people who have shipped model-backed features — owned the evaluation metrics, made the ship-or-hold call on model quality, designed for non-deterministic user experience — and people who added a chatbot tab to something and updated their title. Shipping a chatbot is not nothing. It is also not evidence of the skill that matters, and the application will not tell the two apart on its own.

The red flags cluster tightly, and once you know them they are hard to unsee:

  • The model was 'accurate' — with no definition of what accurate meant for the task, how it was sampled, or how a regression would have been caught. Real owners have numbers and a method; rebranders have adjectives.
  • Every problem is a model problem. Someone who reaches for a model reflexively, and can't name a case where a plain rule would serve users better, has not done the judgement half of the job.
  • The demo is the story. A slick prototype with no account of the unhappy path — what happens when the model is wrong — describes the easy 10% of the work.
  • No failure they'll talk about. Anyone who owned a real model-backed feature owned a real model failure. A candidate with only wins either wasn't close to the work or isn't being straight with you.
  • Roadmap by cherry-picked demo, not by honest limits — a tell that they've never had to reason about error rates with a skeptical stakeholder.

The most common mis-hire is the confident demo-driver: the candidate who dazzles with a slick prototype and can't tell you when the feature should refuse to answer. A demo shows the happy path. The job is the unhappy path. Screen for the second one, or you will hire for the first — and feel it a quarter later.

There is also an honest possibility worth naming in your own head before you post the role: you may not need an AI product manager yet. Wiring one model call behind an existing feature does not require a dedicated hire — your current product manager and one capable engineer can own it. You need this role when the probabilistic component becomes central to the product's value, when 'is the output good enough to ship' keeps recurring as a judgement call, and when a confidently wrong model costs enough that someone must own that failure mode full time. If you're unsure which you have, that uncertainty is itself the answer: start with your existing product manager, and hire the specialist when the judgement calls start piling up. Reaching for one too early usually ends with an expensive person managing a feature that didn't need them.

Once you are sure you need the role, the job description is your first honest filter. Write the responsibilities as behaviours, keep the AI fluency expectations concrete, and drop the credential lines that predict nothing. Then assess against it. If you are staffing the adjacent parts of this shift, the AI governance lead job description is a useful sibling — the two roles increasingly hand work to each other. And prompt engineering for product managers covers the prototyping literacy this role now assumes. A job description that reads like a spec you can test against is the whole ambition — it is, not coincidentally, how we think about the work itself.

Job description templatesAI product managerAI-era rolesProduct managementTechnical hiring
A

Written by

Aayesha Patel · Co-founder, Hanzomon Inc

Co-founder of Hanzomon. Writes about skills-based hiring, fair assessment and building a better candidate experience.

Frequently asked questions

What should an AI product manager job description include?

The usual four blocks — about the role, responsibilities, requirements, nice to have — plus two most templates skip. First, an AI fluency expectations section that spells out writing evaluation criteria for model output, prompt literacy for prototyping, and the judgement to say when a workflow should not use a model. Second, responsibilities framed around holding a quality bar for non-deterministic output, not shipping fixed specs.

Is an AI product manager different from a product manager?

Yes, in one specific way. A product manager ships features that behave the same every time, so success is a clear spec. An AI product manager ships features whose output varies and is sometimes confidently wrong, so the work shifts to defining what good output means, reasoning about error rates, and designing for failure. Same core craft — discovery, prioritisation, stakeholder alignment — built on a component you cannot fully specify in advance.

What skills should an AI product manager have?

Evaluation literacy first: can they define what good output means for a feature and measure it? Then judgement about when not to use a model, comfort explaining error rates to non-technical stakeholders, and the design sense to build guardrails for when the model is wrong. Deep machine-learning maths is optional. The ability to scope a probabilistic problem and hold a quality bar under pressure is not.

Do you need degree requirements in an AI product manager job description?

No. A degree line filters out capable people and predicts almost nothing about whether someone can hold a quality bar on non-deterministic output. The same goes for fixed years-with-a-named-tool bars — the tooling moves faster than any tenure clock. Describe the observable behaviours the role needs and assess for them directly. Skills-based criteria widen your pool and predict performance better than credentials.

How do you write the AI fluency section of a job description?

Name concrete, checkable behaviours rather than a tools wish-list. For an AI product manager, that means writing evaluation criteria for model outputs, enough prompt literacy to prototype and pressure-test a feature, reasoning about error rates aloud, and judging when a workflow should not use AI at all. Write bullets an applicant could recognise themselves in and an interviewer could test against. Avoid listing model names that will date within a quarter.

Related posts

See it on your own job description

Join the early-access waitlist and watch H-Evaluate build an assessment for a real role.

See it on your own job description