All posts

Hiring · July 21, 2026 · 9 min read

Prompt engineering by role: what good looks like across your job verticals

Prompt engineering is now a core skill in almost every role — but 'good' looks completely different for an engineer, an analyst, a support agent or an SDR. A practical, per-vertical guide with real use cases.

By Jakir Patel · Founder, Hanzomon

Share

Part of The five pillars of a hire: what great assessments actually measure

Hiring
On this page

A year ago, 'prompt engineering' sounded like a niche specialty. Now it's a baseline skill in almost every role that touches a keyboard — and it's one of the clearest things the AI Sandbox surfaces. But here's what most hiring teams miss: good prompting doesn't look the same across jobs. What separates a strong engineer isn't what separates a strong support agent. This is a practical, per-vertical guide to what 'good' actually looks like.

Prompt engineering in practice: the AI Sandbox puts a candidate in a role-relevant task with AI tools and assesses how they direct, verify and correct.

Prompt engineering isn't one skill — it's role-shaped

The underlying competence is universal: give the tool clear context and constraints, then verify and correct what comes back. It's part of AI fluency, and the strongest signal is always the same — catching the AI when it's confidently wrong. But the task and the failure modes are role-specific, which is why phrasing tricks don't predict much and a realistic, role-relevant task does. Here's how it plays out across the verticals we assess.

Software engineer

Ask an assistant to write an async API client with retry logic. A strong candidate specifies the constraints up front (modern async patterns, explicit error handling on the right exceptions), reads the generated code, and catches the subtle problems — a deprecated call, a retry that swallows the wrong error, a missing edge case — then fixes and tests them. Good prompting here is inseparable from code review. See what else to test in an engineering assessment.

  • Good: constrains the request, reads the output, catches deprecated or unsafe code, verifies with a quick test.
  • Weak: pastes the generated function and ships it, unhandled edge cases and all.

Data analyst

Ask for a SQL query — say, net revenue by month excluding internal test accounts — or for an interpretation of a result. A strong candidate defines the metric and the exclusions in the prompt, then sanity-checks the numbers against what they know, and questions a confident-but-misleading aggregate rather than pasting it into a report. The prompting skill and the analytical judgment are the same muscle. More on the data analyst assessment.

  • Good: specifies the metric and exclusions, sanity-checks the output, distrusts a number the data doesn't support.
  • Weak: accepts a plausible query, reports a figure that quietly doesn't add up.

Customer support

Ask the AI to draft a reply to a frustrated customer. A strong candidate gives it the relevant policy and the tone they want, then checks the draft for accuracy, softens anything that reads as robotic or dismissive, and removes an over-promise the model slipped in. In support, the prompt is only half the job — the edit is where the judgment shows. See the customer support assessment.

  • Good: supplies policy + tone, verifies accuracy, adjusts to genuine empathy, cuts over-promises.
  • Weak: sends a generic AI reply that's wrong on policy or tonally off.

Sales development

Ask for a cold email to a specific persona. A strong candidate feeds the tool real context (who the buyer is, the value proposition, the one clear ask), personalizes the result, trims it to something a busy VP would actually read, and catches any hallucinated 'fact' about the prospect before it goes out. The prompt sets it up; the judgment keeps it credible. More in the sales development assessment.

  • Good: gives context, personalizes, tightens to a crisp ask, catches a fabricated detail.
  • Weak: fires off a generic templated blast — sometimes with a made-up fact in it.

Product manager

Ask for a first draft of a spec or a prioritization. A strong PM frames the problem and constraints clearly, uses the AI to get a fast starting point, and then applies real product judgment — catching a flawed assumption, cutting scope the model over-added, grounding it in users and metrics rather than shipping the draft as-is. See the product manager assessment.

  • Good: frames the problem, uses AI for a draft, then applies judgment — catches bad assumptions, ties it to real goals.
  • Weak: ships an AI-generated spec with no product thinking layered on top.

Notice the through-line: across every vertical, the differentiator isn't the phrasing of the prompt — it's whether the candidate catches the AI when it's wrong. Verification, not generation, is the skill that separates strong from weak.

Go deeper: a full guide for each role

Each vertical gets its own deep-dive — the best practices, a worked prompt, the good-versus-weak signals and how we assess it:

How to assess it

You can't measure any of this with a quiz about prompts, and you certainly can't measure it by banning AI. You measure it by putting the candidate in a realistic, role-relevant task with AI tools available and watching how they work — which is exactly what an AI Sandbox assessment does, and how AI fluency is scored as a pillar. It's the honest way to test the job as it's actually done — the core idea of AI-native hiring. You can watch a role-tuned assessment get composed to see where prompting fits per vertical.

Prompt engineering isn't a phrasing trick you hire for. It's judgment under an AI's confident wrongness — and it looks different in every seat at the table.
Prompt engineeringAI fluencyAI SandboxAssessment design
J

Written by

Jakir Patel · Founder, Hanzomon

Building H-Evaluate — AI-native, quality-gated hiring assessments. Writes about assessment engineering, hiring integrity and compliance-first AI.

Frequently asked questions

Is prompt engineering a real hiring skill?

Yes — for any role that now works with AI tools, how well someone directs, verifies and corrects an AI is a direct read on their output quality and their risk of shipping the tool's mistakes. It's less about clever phrasing and more about judgment: giving the right context and constraints, then checking what comes back.

Does good prompting look the same across roles?

No. The underlying competence is universal — direct the tool clearly, verify its output, correct course — but the task and the failure modes are role-specific. A strong engineer prompt catches a deprecated API call; a strong support prompt fixes tone and a policy error; a strong analyst prompt questions a misleading number. That's why it's best assessed on a role-relevant task.

How do you assess prompt engineering?

Put the candidate in a realistic, role-relevant task with AI tools available and watch how they work — the AI Sandbox. You're grading the collaboration: how they frame the request, whether they catch the AI when it's wrong, and whether the finished work clears the bar. Phrasing tricks don't survive that; judgment does.

Related posts

See it on your own job description

Join the early-access waitlist and watch H-Evaluate build an assessment for a real role.

See it on your own job description