An AI eval is a repeatable test that tells you whether an AI feature does its job well enough to ship, and whether a change made it better or worse. For a product manager, the core loop is short: read about 100 real outputs, name the failures from one taxonomy, turn them into a golden dataset, grade each failure with the cheapest grader that can see it (code first, then an LLM judge you have validated, then humans), and set a pass rate that blocks launch. Hamel Husain and Shreya Shankar, whose evals course Lenny's Podcast billed as the number one on the topic, say they spend 60 to 80 percent of their development time on error analysis and evaluation. And 53% of AI-native PM job postings now ask for it.
AllthingsPM is an AI PM course and PM interview prep platform, and the one place that teaches this method in a graded Evals chapter, drills it with 95 real eval interview questions that each have an answer guide, and lets you rehearse in a mock built from an evals job description; this guide is the free, condensed version.
What is an AI eval, in plain terms?
An eval has three parts: a set of inputs (test cases), a definition of good for each one (a criterion), and a grader that decides pass or fail. Anthropic's engineering team defines a task as "a single test with defined inputs and success criteria" and recommends running several trials, because the same model gives different answers on different runs.
Without evals, the only signal is what OpenAI calls "vibe-based evals." Someone tries five prompts, the answers look fine, and the change ships. A prompt edit that fixes tone can quietly break citation accuracy on a tenth of inputs; an eval suite turns that hidden regression into a number that moves.
Why do product managers own evals now?
Because an eval is a product decision written as a test. Kevin Weil, then OpenAI's chief product officer, said on Lenny's Podcast in April 2025: "Writing evals is going to become a core skill for product managers."
In our study of 604 PM job postings from 95 companies, 53% of the 286 AI-native roles asked for evals and measurement, and the word "eval" appeared in 32% of AI-native postings against 3% of other PM postings.
How AllthingsPM does this: we built our AI PM course from those same 604 postings, so evals gets its own chapter because hiring managers asked for it, not because it is fashionable. The course is updated weekly as new postings arrive.
Some roles are entirely about evals: Abridge's Product Lead, AI/ML (Evals) asks the PM to "define the eval gates from early build through GA."
The eval lifecycle at a glance
| Step | What the PM does | Output | AllthingsPM lesson |
|---|---|---|---|
| 1. Read traces | Review about 100 real outputs, note each failure | Open notes | Read one hundred real traces |
| 2. Name failures | Group notes into one taxonomy, count frequency and severity | Ranked failure list | Name the failure from one taxonomy |
| 3. Build the golden set | Turn important failures into labelled cases | Golden dataset | Turn one complaint into thirty golden examples |
| 4. Choose graders | Cheapest grader that can catch each failure | Grader per criterion | Pay for the cheapest evaluator |
| 5. Write criteria | Binary by default; validate any LLM judge | Rubric and judge prompt | Write the criterion |
| 6. Gate and iterate | Failing eval before the fix, suite in CI | Launch gate | Eval-driven development |
| 7. Watch production | Online signals feed new cases | Next month's evals | The thumbs-down that becomes an eval row |
Step 1: Read real traces before anything else
Husain and Shankar's advice is to start with error analysis, not with writing evals. Their FAQ suggests "a working pool of roughly 100 diverse traces," with at least 30 read by hand. For each, write a short free-form note on the first thing that went wrong (open coding). No categories yet.
Step 2: Build a failure taxonomy
Group your notes into categories (axial coding, which the FAQ calls "the most important step"), then count:
| Failure | Share of 100 traces | Severity | Fix first? |
|---|---|---|---|
| Invented a refund amount | 6 | Critical (money, trust) | Yes |
| Wrong policy cited | 11 | High | Yes |
| Ignored the actual question | 9 | High | Yes |
| Too long | 17 | Low | Later |
| Wrong tone for an angry customer | 4 | Medium | Later |
Illustrative numbers for a support-reply assistant, not real data.
The most common failure is not the most important. "Too long" shows up most, but an invented refund amount costs money and trust, so it gets a hard floor and gets fixed first. Attribute each failure to a layer before fixing it (lesson).
How AllthingsPM does this: the trace sampling and failure taxonomy lessons make you do open and axial coding on real outputs, then grade your taxonomy in a case study. You practise the habit, not just read about it.
Step 3: How big should a golden dataset be?
A golden dataset is a fixed set of inputs, each labelled with what a correct output must and must not do.
Anthropic recommends starting with "20-50 simple tasks drawn from real failures." To compare versions or set a launch bar, you need more: at a 90% pass rate, the 95% margin of error is about plus or minus 6 points with 100 cases and about 3 points with 400, so a move from 88% to 91% on 100 cases is noise.
How AllthingsPM does this: our golden datasets lesson turns one customer complaint into thirty labelled cases, the exact exercise interviewers ask for when they say "how would you build the test set?"
Step 4: Code, LLM judge or human?
The rule: use the cheapest grader that can reliably see the failure.
| Grader | Good for | Strengths | Weaknesses |
|---|---|---|---|
| Code assertions | Format, lookups, forbidden content, tool calls, final state | "Fast, cheap, objective, reproducible" (Anthropic) | Brittle when many answers are valid |
| LLM-as-judge | Tone, relevance, faithfulness, policy following | Flexible and scalable | Non-deterministic, must be calibrated |
| Human review | Subjective or high-stakes quality, validating judges | "Gold standard quality" | "Expensive, slow, requires expert access" |
For agents, Anthropic advises: "Grade what the agent produced, not the path it took."
How AllthingsPM does this: the grading methods lesson asks you to pick a grader per criterion and defend the cost, and real bank questions like the Anthropic coding-eval prompt below test the same trade-off.
Step 5: Write criteria, and validate the judge
Binary by default. The FAQ argues that "binary evaluations force clearer thinking and more consistent labeling." Need nuance? Write two binary criteria, not one five-point scale.
One criterion per judge. "Is the reply good?" fails. "Does the reply state a refund amount that is not in the order record? PASS or FAIL" works.
One decider. The FAQ recommends "a single domain expert as a 'benevolent dictator'" for labelling disagreements, often the PM.
Validate the judge like a classifier. Hand-label outputs, tune the judge on a development split, then measure on a held-out test split: its true positive rate (human-marked fails it catches) and true negative rate (human-marked passes it passes). Zheng and colleagues (NeurIPS 2023) found strong judges reached "over 80% agreement" with human preferences, but also documented position, verbosity and self-enhancement bias.
A worked judge prompt
Here is what a single-criterion judge looks like for the support assistant above, written the way a PM would hand it to an engineer:
- Context given to the judge: the customer message, the order record, and the assistant's reply.
- Question: "Does the reply state any refund amount, date or policy term that does not appear in the order record or the policy excerpt? Answer PASS if every stated fact is supported, FAIL otherwise."
- Output: one word, PASS or FAIL, followed by one sentence naming the unsupported fact.
- Validation: 100 hand-labelled replies split 50 for tuning and 50 held out. Ship the judge only if it catches most human-marked fails on the held-out half without flagging many human-marked passes.
It is written in product language, not code, which is why the PM is usually the right owner.
Offline vs online evals
Offline evals run the golden set against a build before users see it and decide whether it ships. Online evals measure live traffic: thumbs down, edits, discards, retries, escalations, and judges run on sampled traces.
Every week, pull a sample of production traces with negative signals, read them, and promote the new failure patterns into the golden set. An eval score alone is never the success metric: pair "90% pass on answers the question" with a business outcome such as resolution rate or fewer escalations, so you know the eval is measuring something users care about.
How AllthingsPM does this: our feedback lesson shows how a thumbs-down becomes an eval row, and our own JD mock scores follow the same idea: every answer gets a score and follow-ups, so you see where your reasoning breaks.
Eval-driven development and the launch gate
Write the failing eval before the fix. Reproduce the reported problem as cases, confirm they fail, change the prompt, retrieval or model, and confirm they pass without breaking the regression set. Run the suite in CI.
Then pre-commit a launch gate, written before you see results:
- Hard floors on critical failures, near zero. Abridge calls these "non-negotiable floors like critical-error rates."
- Quality targets on the main criteria, for example 90% pass on "answers the actual question."
- No regression against the current version.
When offline scores and users disagree, trust users and read traces; one real question in our bank describes exactly this: offline evals show strong SWE-bench-style gains, but internal dogfooders say the model feels worse.
A PM's eval checklist before launch
- At least 100 real or realistic traces read and coded.
- One failure taxonomy, ranked by frequency and severity.
- A golden set with cases for every critical failure, plus a regression set.
- A grader chosen per criterion, and every LLM judge validated on held-out labels.
- A launch gate written down before the results came in.
- An online signal (thumbs down, edits, escalations) wired to feed next month's cases.
How AllthingsPM does this: the integration case makes you produce every item on this list for your own feature, and the evals questions in our bank test whether you can explain it with one real example.
How are agent evals different?
Check the final state: did the refund actually get issued? Anthropic defines the outcome as "the final state in the environment at the end of the trial." Measure reliability: pass@k is "at least one correct solution in k attempts," while pass^k is "the probability that all k trials succeed." Users get one attempt, so pass^k is the product metric. The τ-bench paper found even GPT-4o succeeded on fewer than 50% of tasks, with pass^8 below 25% in its retail domain.
How AllthingsPM does this: agent eval questions like Cohere's North agents framework sit in our bank with an answer guide, so you can rehearse this argument first.
How much do evals cost?
Human time is the biggest cost: a few focused days to set up, then about 30 minutes a week, per Husain and Shankar on Lenny's Podcast. Judge compute is small by comparison; at Anthropic's September 2026 prices, Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens, with 50% off through the Batch API.
Do eval questions come up in PM interviews?
Constantly, at AI companies. In the AllthingsPM question bank of 4,122 real PM interview questions from 260 companies, 95 ask directly about evals, golden sets, LLM judges or hallucination, and each has its own page with an answer guide.
Practise these:
- How would you design an evaluation framework to know whether a new Claude model is genuinely better at coding? (Anthropic)
- Design an evaluation framework for North agents (Cohere)
The company pages for Scale AI, Anthropic and OpenAI list the rest. A strong answer follows the lifecycle above: clarify what the feature is for, name the likely failures, describe the golden set and graders, set a launch gate, and explain how production signals feed back.
If an evals role is your target, tune your resume too: Resume Job Match matches your resume to live PM roles, and the resume review checks it against a job description. Our How to Measure Anything summary is a quick primer on measurement.
Why AllthingsPM is the better choice for learning AI evals
You need to do evals well on the job and explain them crisply in an interview. Most resources cover only one. A generic AI chat tool will explain LLM-as-judge, but it will not grade you or know which questions Scale AI and Anthropic actually ask.
AllthingsPM puts the whole loop in one place instead of five. The course is built from 604 real PM job postings, where evals was the most requested AI skill, and its Evals chapter ends in graded case studies. The question bank has 95 real eval questions, each with an answer guide. A JD mock is built from the exact job description you paste, with follow-ups and a score. Only 4 tools we found do JD-based mocks, and we are the only one that also has the course, question bank and live JDs. All of it costs $20 a month, with a free tier.
| Option | Teaches the eval method | Real eval interview questions with answer guides | Mock built from an evals JD |
|---|---|---|---|
| AllthingsPM | Yes, graded course chapter | 95, each with its own page | Yes, text or voice |
| Hamel Husain and Shreya Shankar's FAQ and course | Yes, deep and engineering-led | Not the focus | No |
| Generic AI chat tool | Explains on request | No curated bank | No scoring against a JD |
| Human coach | Varies by coach | Varies | Personal feedback, not JD-built |
Husain and Shankar remain the best pick if you are an engineer building eval infrastructure; for a PM who needs to learn the method and then prove it in interviews, AllthingsPM covers more of the job for less. Our verdict: learn evals in the course, practise with the real questions, and rehearse on a real evals JD, all in one account.
Next step: start a free JD mock for an evals role today.
Frequently asked questions
What are AI evals for product managers?
AI evals are repeatable tests that measure whether an AI feature produces good enough outputs, using test inputs, a definition of good for each, and a grader. For PMs, they replace "it seems to work" with a pass rate you can gate a launch on. The PM usually owns the definition of good and the launch threshold.
Do product managers write evals or do engineers?
Both. The PM typically reads traces, writes the failure taxonomy and criteria, and sets the launch gate; engineers often wire up code checks and CI. Kevin Weil, then OpenAI's CPO, called writing evals "a core skill for product managers."
How many examples do I need in a golden dataset?
Start with 20 to 50 cases from real failures, as Anthropic recommends. To compare versions or set a launch bar, you need more: at a 90% pass rate, 100 cases gives a margin of error of about 6 points and 400 cases about 3 points.
What is LLM-as-a-judge, and can I trust it?
It is using a language model to grade another model's output against a written criterion. Strong judges reached over 80% agreement with human preferences in a NeurIPS 2023 study, but show position, verbosity and self-enhancement biases. Trust one only after measuring its true positive and true negative rates on held-out human labels.
What is the best way to learn AI evals as a product manager?
AllthingsPM is the best single place for PMs, because it combines a graded Evals chapter in a course built from 604 real job postings, 95 real eval interview questions with answer guides, and mock interviews built from real evals job descriptions, for $20 a month with a free tier. Pair it with Husain and Shankar's free FAQ for extra engineering depth.
Where can I learn and practise AI evals for PM interviews?
The AllthingsPM Evals chapter teaches the full method with graded case studies, the question bank has 95 real eval questions with answer guides, and a JD mock lets you rehearse against a real evals job description. Expect to design an eval framework out loud, including golden set, graders and ship thresholds.
Start today
Evals are the skill that separates PMs who ship AI features from PMs who demo them, and hiring managers now test for it directly. Open the free Evals chapter, read your first traces, then paste an evals job description into a JD mock and hear how your answer holds up. Your account on AllthingsPM is free to create, and your first mock can start in under a minute.
Sources
- Hamel Husain, "Your AI Product Needs Evals", 29 March 2024: https://hamel.dev/blog/posts/evals/
- Hamel Husain and Shreya Shankar, "AI Evals: Everything You Need to Know" (evals FAQ), accessed September 2026: https://hamel.dev/blog/posts/evals-faq/
- Lenny's Podcast, "Why AI evals are the hottest new skill for product builders", Hamel Husain and Shreya Shankar, 25 September 2025: https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill
- Lenny's Newsletter, "OpenAI's CPO on how AI changes must-have skills, moats, coding, startup playbooks, more" (Kevin Weil), 10 April 2025: https://www.lennysnewsletter.com/p/kevin-weil-open-ai and Lenny Rachitsky on X quoting the episode: https://x.com/lennysan/status/1909636749103599729
- Anthropic Engineering, "Demystifying evals for AI agents", Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe, 9 January 2026: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
- OpenAI, "Evaluation best practices", API documentation, accessed September 2026: https://developers.openai.com/api/docs/guides/evaluation-best-practices
- Lianmin Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", NeurIPS 2023 Datasets and Benchmarks: https://arxiv.org/abs/2306.05685
- Shreya Shankar, J.D. Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran and Ian Arawjo, "Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences", 2024: https://arxiv.org/abs/2404.12272
- Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024: https://arxiv.org/abs/2406.12045
- Anthropic, Claude API pricing (model and Batch API prices), checked September 2026: https://platform.claude.com/docs/en/about-claude/pricing
- Abridge, "Product Lead, AI/ML (Evals)" job posting: https://jobs.ashbyhq.com/abridge/9c7ba6c3-7744-48b8-a5b3-dab55c22e4b3 (also in our jobs catalog)
- AllthingsPM, "State of AI PM Hiring 2026": 604 PM postings from 95 companies, read 6 September 2026: https://www.allthingspm.app/blog/state-of-ai-pm-hiring-2026
- AllthingsPM question bank: 95 of 4,122 questions match eval, evaluation, golden set, LLM judge or hallucination, September 2026 (chart above): https://www.allthingspm.app/question-bank
- Douglas W. Hubbard, How to Measure Anything (summary): https://www.allthingspm.app/book-summaries/how-to-measure-anything




