Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

AI Evals for Product Managers: The Complete Guide

AI evals are the tests that decide whether an AI feature is good enough to ship. This guide gives PMs the full method, and AllthingsPM lets you learn it in a graded course chapter and practise it in mock interviews built from real evals job descriptions.

AllthingsPM·September 26, 2026·18 min read
A product manager at a long table sorts a tall stack of printed pages into labelled trays while a laptop sits open beside them
Most eval work is reading real outputs and sorting the failures. The tooling comes later.

An AI eval is a repeatable test that tells you whether an AI feature does its job well enough to ship, and whether a change made it better or worse. For a product manager, the core loop is short: read about 100 real outputs, name the failures from one taxonomy, turn them into a golden dataset, grade each failure with the cheapest grader that can see it (code first, then an LLM judge you have validated, then humans), and set a pass rate that blocks launch. Hamel Husain and Shreya Shankar, whose evals course Lenny's Podcast billed as the number one on the topic, say they spend 60 to 80 percent of their development time on error analysis and evaluation. And 53% of AI-native PM job postings now ask for it.

AllthingsPM is an AI PM course and PM interview prep platform, and the one place that teaches this method in a graded Evals chapter, drills it with 95 real eval interview questions that each have an answer guide, and lets you rehearse in a mock built from an evals job description; this guide is the free, condensed version.

What is an AI eval, in plain terms?

An eval has three parts: a set of inputs (test cases), a definition of good for each one (a criterion), and a grader that decides pass or fail. Anthropic's engineering team defines a task as "a single test with defined inputs and success criteria" and recommends running several trials, because the same model gives different answers on different runs.

Without evals, the only signal is what OpenAI calls "vibe-based evals." Someone tries five prompts, the answers look fine, and the change ships. A prompt edit that fixes tone can quietly break citation accuracy on a tenth of inputs; an eval suite turns that hidden regression into a number that moves.

Why do product managers own evals now?

Because an eval is a product decision written as a test. Kevin Weil, then OpenAI's chief product officer, said on Lenny's Podcast in April 2025: "Writing evals is going to become a core skill for product managers."

In our study of 604 PM job postings from 95 companies, 53% of the 286 AI-native roles asked for evals and measurement, and the word "eval" appeared in 32% of AI-native postings against 3% of other PM postings.

How AllthingsPM does this: we built our AI PM course from those same 604 postings, so evals gets its own chapter because hiring managers asked for it, not because it is fashionable. The course is updated weekly as new postings arrive.

Some roles are entirely about evals: Abridge's Product Lead, AI/ML (Evals) asks the PM to "define the eval gates from early build through GA."

The eval lifecycle at a glance

StepWhat the PM doesOutputAllthingsPM lesson
1. Read tracesReview about 100 real outputs, note each failureOpen notesRead one hundred real traces
2. Name failuresGroup notes into one taxonomy, count frequency and severityRanked failure listName the failure from one taxonomy
3. Build the golden setTurn important failures into labelled casesGolden datasetTurn one complaint into thirty golden examples
4. Choose gradersCheapest grader that can catch each failureGrader per criterionPay for the cheapest evaluator
5. Write criteriaBinary by default; validate any LLM judgeRubric and judge promptWrite the criterion
6. Gate and iterateFailing eval before the fix, suite in CILaunch gateEval-driven development
7. Watch productionOnline signals feed new casesNext month's evalsThe thumbs-down that becomes an eval row

Step 1: Read real traces before anything else

Husain and Shankar's advice is to start with error analysis, not with writing evals. Their FAQ suggests "a working pool of roughly 100 diverse traces," with at least 30 read by hand. For each, write a short free-form note on the first thing that went wrong (open coding). No categories yet.

Step 2: Build a failure taxonomy

Group your notes into categories (axial coding, which the FAQ calls "the most important step"), then count:

FailureShare of 100 tracesSeverityFix first?
Invented a refund amount6Critical (money, trust)Yes
Wrong policy cited11HighYes
Ignored the actual question9HighYes
Too long17LowLater
Wrong tone for an angry customer4MediumLater

Illustrative numbers for a support-reply assistant, not real data.

The most common failure is not the most important. "Too long" shows up most, but an invented refund amount costs money and trust, so it gets a hard floor and gets fixed first. Attribute each failure to a layer before fixing it (lesson).

How AllthingsPM does this: the trace sampling and failure taxonomy lessons make you do open and axial coding on real outputs, then grade your taxonomy in a case study. You practise the habit, not just read about it.

Step 3: How big should a golden dataset be?

A golden dataset is a fixed set of inputs, each labelled with what a correct output must and must not do.

Anthropic recommends starting with "20-50 simple tasks drawn from real failures." To compare versions or set a launch bar, you need more: at a 90% pass rate, the 95% margin of error is about plus or minus 6 points with 100 cases and about 3 points with 400, so a move from 88% to 91% on 100 cases is noise.

How AllthingsPM does this: our golden datasets lesson turns one customer complaint into thirty labelled cases, the exact exercise interviewers ask for when they say "how would you build the test set?"

Step 4: Code, LLM judge or human?

The rule: use the cheapest grader that can reliably see the failure.

GraderGood forStrengthsWeaknesses
Code assertionsFormat, lookups, forbidden content, tool calls, final state"Fast, cheap, objective, reproducible" (Anthropic)Brittle when many answers are valid
LLM-as-judgeTone, relevance, faithfulness, policy followingFlexible and scalableNon-deterministic, must be calibrated
Human reviewSubjective or high-stakes quality, validating judges"Gold standard quality""Expensive, slow, requires expert access"

For agents, Anthropic advises: "Grade what the agent produced, not the path it took."

How AllthingsPM does this: the grading methods lesson asks you to pick a grader per criterion and defend the cost, and real bank questions like the Anthropic coding-eval prompt below test the same trade-off.

Step 5: Write criteria, and validate the judge

Binary by default. The FAQ argues that "binary evaluations force clearer thinking and more consistent labeling." Need nuance? Write two binary criteria, not one five-point scale.

One criterion per judge. "Is the reply good?" fails. "Does the reply state a refund amount that is not in the order record? PASS or FAIL" works.

One decider. The FAQ recommends "a single domain expert as a 'benevolent dictator'" for labelling disagreements, often the PM.

Validate the judge like a classifier. Hand-label outputs, tune the judge on a development split, then measure on a held-out test split: its true positive rate (human-marked fails it catches) and true negative rate (human-marked passes it passes). Zheng and colleagues (NeurIPS 2023) found strong judges reached "over 80% agreement" with human preferences, but also documented position, verbosity and self-enhancement bias.

A worked judge prompt

Here is what a single-criterion judge looks like for the support assistant above, written the way a PM would hand it to an engineer:

  1. Context given to the judge: the customer message, the order record, and the assistant's reply.
  2. Question: "Does the reply state any refund amount, date or policy term that does not appear in the order record or the policy excerpt? Answer PASS if every stated fact is supported, FAIL otherwise."
  3. Output: one word, PASS or FAIL, followed by one sentence naming the unsupported fact.
  4. Validation: 100 hand-labelled replies split 50 for tuning and 50 held out. Ship the judge only if it catches most human-marked fails on the held-out half without flagging many human-marked passes.

It is written in product language, not code, which is why the PM is usually the right owner.

Offline vs online evals

Offline evals run the golden set against a build before users see it and decide whether it ships. Online evals measure live traffic: thumbs down, edits, discards, retries, escalations, and judges run on sampled traces.

Every week, pull a sample of production traces with negative signals, read them, and promote the new failure patterns into the golden set. An eval score alone is never the success metric: pair "90% pass on answers the question" with a business outcome such as resolution rate or fewer escalations, so you know the eval is measuring something users care about.

How AllthingsPM does this: our feedback lesson shows how a thumbs-down becomes an eval row, and our own JD mock scores follow the same idea: every answer gets a score and follow-ups, so you see where your reasoning breaks.

Eval-driven development and the launch gate

Write the failing eval before the fix. Reproduce the reported problem as cases, confirm they fail, change the prompt, retrieval or model, and confirm they pass without breaking the regression set. Run the suite in CI.

Then pre-commit a launch gate, written before you see results:

  • Hard floors on critical failures, near zero. Abridge calls these "non-negotiable floors like critical-error rates."
  • Quality targets on the main criteria, for example 90% pass on "answers the actual question."
  • No regression against the current version.

When offline scores and users disagree, trust users and read traces; one real question in our bank describes exactly this: offline evals show strong SWE-bench-style gains, but internal dogfooders say the model feels worse.

A PM's eval checklist before launch

  • At least 100 real or realistic traces read and coded.
  • One failure taxonomy, ranked by frequency and severity.
  • A golden set with cases for every critical failure, plus a regression set.
  • A grader chosen per criterion, and every LLM judge validated on held-out labels.
  • A launch gate written down before the results came in.
  • An online signal (thumbs down, edits, escalations) wired to feed next month's cases.

How AllthingsPM does this: the integration case makes you produce every item on this list for your own feature, and the evals questions in our bank test whether you can explain it with one real example.

How are agent evals different?

Check the final state: did the refund actually get issued? Anthropic defines the outcome as "the final state in the environment at the end of the trial." Measure reliability: pass@k is "at least one correct solution in k attempts," while pass^k is "the probability that all k trials succeed." Users get one attempt, so pass^k is the product metric. The τ-bench paper found even GPT-4o succeeded on fewer than 50% of tasks, with pass^8 below 25% in its retail domain.

How AllthingsPM does this: agent eval questions like Cohere's North agents framework sit in our bank with an answer guide, so you can rehearse this argument first.

How much do evals cost?

Human time is the biggest cost: a few focused days to set up, then about 30 minutes a week, per Husain and Shankar on Lenny's Podcast. Judge compute is small by comparison; at Anthropic's September 2026 prices, Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens, with 50% off through the Batch API.

Do eval questions come up in PM interviews?

Constantly, at AI companies. In the AllthingsPM question bank of 4,122 real PM interview questions from 260 companies, 95 ask directly about evals, golden sets, LLM judges or hallucination, and each has its own page with an answer guide.

Bar chart led by an AllthingsPM (us) row: 95 eval questions to practise in the AllthingsPM question bank, then by company: Scale AI 19, Anthropic 16, Sierra 13, OpenAI 11, Abridge 7, Decagon 5, Ramp 4, Figma 3, Suno 3
Source: AllthingsPM question bank, 4,122 questions, September 2026. Matches on eval, evaluation, golden set, LLM judge and hallucination

Practise these:

The company pages for Scale AI, Anthropic and OpenAI list the rest. A strong answer follows the lifecycle above: clarify what the feature is for, name the likely failures, describe the golden set and graders, set a launch gate, and explain how production signals feed back.

If an evals role is your target, tune your resume too: Resume Job Match matches your resume to live PM roles, and the resume review checks it against a job description. Our How to Measure Anything summary is a quick primer on measurement.

Why AllthingsPM is the better choice for learning AI evals

You need to do evals well on the job and explain them crisply in an interview. Most resources cover only one. A generic AI chat tool will explain LLM-as-judge, but it will not grade you or know which questions Scale AI and Anthropic actually ask.

AllthingsPM puts the whole loop in one place instead of five. The course is built from 604 real PM job postings, where evals was the most requested AI skill, and its Evals chapter ends in graded case studies. The question bank has 95 real eval questions, each with an answer guide. A JD mock is built from the exact job description you paste, with follow-ups and a score. Only 4 tools we found do JD-based mocks, and we are the only one that also has the course, question bank and live JDs. All of it costs $20 a month, with a free tier.

OptionTeaches the eval methodReal eval interview questions with answer guidesMock built from an evals JD
AllthingsPMYes, graded course chapter95, each with its own pageYes, text or voice
Hamel Husain and Shreya Shankar's FAQ and courseYes, deep and engineering-ledNot the focusNo
Generic AI chat toolExplains on requestNo curated bankNo scoring against a JD
Human coachVaries by coachVariesPersonal feedback, not JD-built

Husain and Shankar remain the best pick if you are an engineer building eval infrastructure; for a PM who needs to learn the method and then prove it in interviews, AllthingsPM covers more of the job for less. Our verdict: learn evals in the course, practise with the real questions, and rehearse on a real evals JD, all in one account.

Next step: start a free JD mock for an evals role today.

Frequently asked questions

What are AI evals for product managers?

AI evals are repeatable tests that measure whether an AI feature produces good enough outputs, using test inputs, a definition of good for each, and a grader. For PMs, they replace "it seems to work" with a pass rate you can gate a launch on. The PM usually owns the definition of good and the launch threshold.

Do product managers write evals or do engineers?

Both. The PM typically reads traces, writes the failure taxonomy and criteria, and sets the launch gate; engineers often wire up code checks and CI. Kevin Weil, then OpenAI's CPO, called writing evals "a core skill for product managers."

How many examples do I need in a golden dataset?

Start with 20 to 50 cases from real failures, as Anthropic recommends. To compare versions or set a launch bar, you need more: at a 90% pass rate, 100 cases gives a margin of error of about 6 points and 400 cases about 3 points.

What is LLM-as-a-judge, and can I trust it?

It is using a language model to grade another model's output against a written criterion. Strong judges reached over 80% agreement with human preferences in a NeurIPS 2023 study, but show position, verbosity and self-enhancement biases. Trust one only after measuring its true positive and true negative rates on held-out human labels.

What is the best way to learn AI evals as a product manager?

AllthingsPM is the best single place for PMs, because it combines a graded Evals chapter in a course built from 604 real job postings, 95 real eval interview questions with answer guides, and mock interviews built from real evals job descriptions, for $20 a month with a free tier. Pair it with Husain and Shankar's free FAQ for extra engineering depth.

Where can I learn and practise AI evals for PM interviews?

The AllthingsPM Evals chapter teaches the full method with graded case studies, the question bank has 95 real eval questions with answer guides, and a JD mock lets you rehearse against a real evals job description. Expect to design an eval framework out loud, including golden set, graders and ship thresholds.

Start today

Evals are the skill that separates PMs who ship AI features from PMs who demo them, and hiring managers now test for it directly. Open the free Evals chapter, read your first traces, then paste an evals job description into a JD mock and hear how your answer holds up. Your account on AllthingsPM is free to create, and your first mock can start in under a minute.

Sources

  1. Hamel Husain, "Your AI Product Needs Evals", 29 March 2024: https://hamel.dev/blog/posts/evals/
  2. Hamel Husain and Shreya Shankar, "AI Evals: Everything You Need to Know" (evals FAQ), accessed September 2026: https://hamel.dev/blog/posts/evals-faq/
  3. Lenny's Podcast, "Why AI evals are the hottest new skill for product builders", Hamel Husain and Shreya Shankar, 25 September 2025: https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill
  4. Lenny's Newsletter, "OpenAI's CPO on how AI changes must-have skills, moats, coding, startup playbooks, more" (Kevin Weil), 10 April 2025: https://www.lennysnewsletter.com/p/kevin-weil-open-ai and Lenny Rachitsky on X quoting the episode: https://x.com/lennysan/status/1909636749103599729
  5. Anthropic Engineering, "Demystifying evals for AI agents", Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe, 9 January 2026: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
  6. OpenAI, "Evaluation best practices", API documentation, accessed September 2026: https://developers.openai.com/api/docs/guides/evaluation-best-practices
  7. Lianmin Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", NeurIPS 2023 Datasets and Benchmarks: https://arxiv.org/abs/2306.05685
  8. Shreya Shankar, J.D. Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran and Ian Arawjo, "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences", 2024: https://arxiv.org/abs/2404.12272
  9. Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024: https://arxiv.org/abs/2406.12045
  10. Anthropic, Claude API pricing (model and Batch API prices), checked September 2026: https://platform.claude.com/docs/en/about-claude/pricing
  11. Abridge, "Product Lead, AI/ML (Evals)" job posting: https://jobs.ashbyhq.com/abridge/9c7ba6c3-7744-48b8-a5b3-dab55c22e4b3 (also in our jobs catalog)
  12. AllthingsPM, "State of AI PM Hiring 2026": 604 PM postings from 95 companies, read 6 September 2026: https://www.allthingspm.app/blog/state-of-ai-pm-hiring-2026
  13. AllthingsPM question bank: 95 of 4,122 questions match eval, evaluation, golden set, LLM judge or hallucination, September 2026 (chart above): https://www.allthingspm.app/question-bank
  14. Douglas W. Hubbard, How to Measure Anything (summary): https://www.allthingspm.app/book-summaries/how-to-measure-anything
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free