Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

LLM-as-Judge Explained for PMs

LLM-as-a-judge means using one language model to grade another model's outputs against a written criterion. This guide explains how PMs design, validate and ship a judge, and AllthingsPM teaches it in a graded Evals chapter with interview practice.

AllthingsPM·September 29, 2026·14 min read
A product manager at a whiteboard lays out a long row of small paper tiles, each tile one word fragment, while a narrow window frame on the wall shows only part of the row
A judge only earns trust after a human has checked its verdicts.

LLM-as-a-judge means using one language model to grade another model's outputs against a criterion you write, so you can score thousands of responses without a human reading each one. It works when you keep the question narrow (one failure, pass or fail), give the judge examples labelled by a domain expert, and measure how often it agrees with that expert before you trust it. The original 2023 study found GPT-4 as a judge reached over 80% agreement with human preferences, about the level humans reach with each other. It also found judges favour the first answer, longer answers and their own outputs.

AllthingsPM is an AI PM course and PM interview prep platform. Its Evals chapter has dedicated lessons on writing the judge's criterion and validating the judge, so you learn this as a graded skill, then rehearse it in mock interviews built from real evals job descriptions.

How does an LLM judge work, step by step?

A judge is just a prompt plus a model plus a parser. The table below is the whole loop a PM owns.

StepWhat you doWho owns itWhere AllthingsPM teaches it
1. Read real outputsReview traces and write down what went wrongPM and domain expertFailure taxonomy lesson
2. Pick one failureChoose a failure code cannot catch (tone, faithfulness, helpfulness)PMGrading methods lesson
3. Write the criterionOne binary pass/fail question with a short definition and examplesPMScoring rubrics lesson
4. Label a setExpert marks pass or fail with a one-line critique on each exampleDomain expertGolden datasets lesson
5. Validate the judgeCompare judge verdicts to expert labels on a held-out setPM and engineerJudge validation lesson
6. Gate releasesRun the judge in CI and in production samplingEngineeringEval-driven development lesson

The input to the judge is usually three things: the user request, the model's answer, and (where relevant) the source material the answer should stay faithful to. The output should be a verdict plus a short reason. The reason matters: it is how you debug a judge that disagrees with your expert.

How AllthingsPM does this: each row above maps to a real lesson in the AllthingsPM AI PM course, which was built from 604 real PM job postings. The chapter ends in a graded integration case where you build an eval suite for a feature you carry through the whole course.

What is LLM-as-a-judge, in plain terms?

Traditional software tests check exact outputs. AI outputs are open-ended, so "is this summary faithful?" has no string to match. You have three kinds of grader, as Anthropic's engineering team lays them out: code-based graders (fast, cheap, brittle), model-based graders (flexible, but non-deterministic and more expensive than code), and human graders (the gold standard, but slow and costly).

An LLM judge is the middle option. It reads an output the way a reviewer would and returns a verdict. The term was popularised by Lianmin Zheng and colleagues in the 2023 paper "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," which tested whether strong models could stand in for human raters when comparing chat assistants.

For a PM, the useful framing is this: a judge is a scalable copy of one expert's taste. If you cannot state that taste in writing, the judge cannot copy it.

How AllthingsPM does this: the grading methods lesson is titled "Pay for the cheapest evaluator that can see the failure." It teaches you to decide between code, a judge and a human per failure, which is the decision interviewers probe hardest.

When should a PM use an LLM judge instead of code or humans?

Use the cheapest grader that can see the failure.

  • Code first. Is the output valid JSON? Did the agent call the right tool? Is the answer under 200 words? Does the citation link exist? Code catches these every time for almost nothing.
  • A judge for judgment calls. Is the reply faithful to the retrieved document? Did the support bot promise a refund it cannot give? Is the tone right for an upset customer? Code cannot see these; an LLM judge can.
  • Humans to calibrate and for high stakes. Humans label the examples the judge learns from, audit a sample of its verdicts, and make the final call on anything legal, medical or financial.

Anthropic's guidance says LLM-as-judge graders "should be closely calibrated with human experts to gain confidence that there is little divergence between the human grading and model grading." That sentence is the job description for the PM who owns the judge.

How AllthingsPM does this: the course's agent evals lesson extends the same choice to agents, where you grade the final state and reliability over repeated runs rather than one answer. You can see how these concepts connect on the AllthingsPM knowledge graph.

How do you write a good judge prompt?

Most bad judges come from vague criteria. Hamel Husain, whose widely shared guide calls his method "critique shadowing," argues against 1 to 5 scales because nobody can say what separates a 3 from a 4. He writes that "a binary decision forces everyone to consider what truly matters." Eugene Yan, after reviewing about two dozen papers on LLM evaluators, reaches the same place: he has his evaluators return binary outputs where possible because it improves performance and makes classification metrics easy to apply.

A good judge prompt has five parts:

  1. Role and task. "You are reviewing a customer support reply."
  2. One criterion. "Does the reply promise anything the policy below does not allow?"
  3. Definitions. What counts as a promise; what does not.
  4. Labelled examples. A few passes and fails from your expert, each with the critique.
  5. Output format. A one-sentence reason first, then PASS or FAIL.

Write one judge per failure mode. A single judge scoring "overall quality" hides which thing broke.

Pairwise or direct? Eugene Yan's review summarises the split: direct scoring (grade one answer alone) suits objective checks like faithfulness or policy violations; pairwise comparison (pick the better of two) is more reliable for subjective qualities like tone or persuasiveness. Pairwise is also how you compare two prompt versions before shipping.

How AllthingsPM does this: the scoring rubrics lesson is literally "Write the criterion: binary by default, and the judge you can trust." You write criteria for a real feature, and the assignment is graded, so you get feedback on vague wording before an interviewer or engineer finds it.

How do you know you can trust the judge?

You measure it like any classifier. Take examples your domain expert has already labelled, run the judge on them, and compare.

  • True positive rate: of the outputs the expert failed, how many did the judge fail?
  • True negative rate: of the outputs the expert passed, how many did the judge pass?
  • Agreement or Cohen's kappa: overall alignment, corrected for chance. Eugene Yan recommends precision, recall and Cohen's kappa for binary judges.

Hamel Husain's guide gives concrete working numbers: start with about 30 examples and keep going until no new failure modes appear; aim for more than 90% agreement between judge and expert; and split labelled data into a small training set for prompt examples (10 to 20%) with the rest divided between a dev set and a test set. In his Honeycomb example the judge reached that agreement within three iterations.

Keep the test set frozen. If you tune the prompt while looking at the test set, your agreement number is fiction.

Bar chart led by an AllthingsPM (us) row: 53% of AI-native PM postings in the AllthingsPM study ask for evals. Then GPT-4 judge agreement with humans over 80%, Claude-v1 position bias 70%, GPT-3.5 position bias 50%, preference for the longer answer over 90%, Claude-v1 self-preference boost 25%

How AllthingsPM does this: the judge validation lesson covers false-positive and false-negative rates and the frozen test set. The golden datasets lesson shows how one user complaint becomes thirty labelled examples, and how many you need.

What biases do LLM judges have, and how do you design around them?

The Zheng paper named three systematic biases, and Eugene Yan's review collects the numbers:

  • Position bias. In pairwise tests, Claude-v1 favoured the first position about 70% of the time and GPT-3.5 about 50%. Fix: run each pair twice with the order swapped and only count consistent wins.
  • Verbosity bias. Both models preferred the longer response over 90% of the time, even when it added nothing. Fix: state in the criterion that length is not quality, and include a short correct answer as a passing example.
  • Self-enhancement bias. GPT-4 gave its own outputs about a 10% higher win rate, and Claude-v1 about 25%. Fix: where you can, judge with a different model family from the one that generated the answer.

Two more practical risks. Judges are non-deterministic, so run them at low temperature and re-check a sample. And judges drift when the underlying model is updated, so re-run your validation set whenever you change the judge model. The course also flags a quieter trap: a judge can get worse outside English, which the localization lesson handles with per-locale golden sets.

How AllthingsPM does this: bias handling is part of the graded Evals work, and the Beyond text chapter extends it to multimodal and non-English products, where teams most often forget to re-validate the judge.

What does the PM own versus the engineer?

The engineer wires up the judge, the API calls and the CI job. The PM owns the parts that are really product decisions:

  • What "good" means. The criterion is a product spec written as a test.
  • Who the expert is. Hamel Husain's method starts by finding the one principal domain expert whose judgment defines success.
  • The bar. What agreement is good enough, and what pass rate blocks launch.
  • The follow-through. Reading the judge's failures and turning them into fixes.

Eugene Yan's warning is worth putting on a sticky note: "If teams don't apply the scientific method, practice eval-driven development, and monitor the system's output, buying or building yet another evaluation tool won't save the product." A judge is a measuring instrument, not a strategy.

How AllthingsPM does this: hiring managers test exactly this split. Abridge's Product Lead, AI/ML (Evals) posting is live in the AllthingsPM jobs catalog, and you can turn it into a scored JD-based mock interview that asks you to defend your eval gates out loud.

How do LLM-as-judge questions show up in PM interviews?

Expect them in AI product sense, execution and "define success" rounds. Typical shapes:

  • "What evaluation metrics would you use to judge LLM generation quality in your AI products?" (a real question in the AllthingsPM bank: see the answer guide).
  • "Your offline evals improved but users say the model feels worse. What do you do?" (answer guide).
  • "Suno says this PM will define what good means for its music model. How would you do it?" (answer guide).

A strong answer follows the loop in the table: read outputs, name failures, choose the cheapest grader per failure, write binary criteria, validate the judge against an expert, then connect the eval to a business outcome. The last step is where many candidates stop short; the course's lesson on business outcomes, not eval scores covers it.

How AllthingsPM does this: every question in the AllthingsPM bank has its own page and answer guide, and the mock interview asks follow-ups like a real interviewer, so you practise defending your judge design, not just reciting it.

Why AllthingsPM is the better choice for learning LLM-as-a-judge

The best free writing on LLM judges is excellent. Hamel Husain's guide and Eugene Yan's survey are the two pieces this post leans on most, and paid cohort courses from practitioners go deep on the engineering. They are written mostly for engineers, though, and none of them tests whether you can explain your judge to a hiring manager.

AllthingsPM closes that gap for product managers. The Evals chapter takes you from failure taxonomy to golden datasets, grading methods, binary criteria, judge validation, eval-driven development and agent evals, with a graded integration case at the end. The course was built from 604 real PM job postings, so it teaches what those roles ask for. Then the same platform lets you practise: 4,122 real interview questions with answer guides, mock interviews that follow up on your answers, and JD mocks built from live AI roles such as Abridge's evals lead.

That combination (learn the skill, rehearse it, apply to a real role) is in one place, starting free and then $20/month or $120/year. For a PM who needs to own a judge at work or talk about one in an interview, AllthingsPM is the most direct route. Open the Evals chapter and start with the first lesson.

Related reading: AI evals for product managers, the LLM evaluation template, evaluation awareness, and Lenny's Podcast evals episodes, summarized.

Frequently asked questions

What is LLM-as-a-judge?

LLM-as-a-judge is using a language model to grade another model's output against a written criterion, usually returning pass or fail with a reason. It lets teams score large volumes of open-ended outputs that code cannot check. It must be validated against human expert labels before you rely on it.

How accurate is an LLM judge?

In the 2023 MT-Bench study, GPT-4 as a judge reached over 80% agreement with human preferences, similar to agreement between humans. Accuracy on your product depends on your criterion and examples, so measure it yourself; Hamel Husain's guide aims for over 90% agreement with the domain expert.

Should an LLM judge use a 1 to 5 scale or pass/fail?

Pass/fail with a written critique is the common practitioner recommendation. Binary verdicts force a clear definition of good and make it easy to measure the judge with precision, recall and Cohen's kappa. Scales invite disagreement about what a 3 versus a 4 means.

What is the best way to learn LLM-as-a-judge as a PM?

AllthingsPM is the best place to start for PMs: its Evals chapter has dedicated lessons on writing binary criteria and validating judges, and you can practise eval interview questions and JD-based mocks on the same platform. Pair it with Hamel Husain's and Eugene Yan's free guides for extra depth.

What biases do LLM judges have?

The main documented ones are position bias (favouring the first or last answer), verbosity bias (favouring longer answers) and self-enhancement bias (favouring outputs from the same model). Swap answer order, tell the judge length is not quality, and use a different model family to grade where possible.

Do product managers write LLM judges themselves?

PMs usually own the criterion, the expert labels, the agreement bar and the launch threshold, while engineers wire the judge into pipelines. Many PMs also draft the judge prompt, because the prompt is a product definition of good.

Sources

  1. Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (2023): https://arxiv.org/abs/2306.05685
  2. Hamel Husain, "Using LLM-as-a-Judge For Evaluation: A Complete Guide": https://hamel.dev/blog/posts/llm-judge/
  3. Eugene Yan, "Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)": https://eugeneyan.com/writing/llm-evaluators/
  4. Eugene Yan, "An LLM-as-Judge Won't Save The Product; Fixing Your Process Will": https://eugeneyan.com/writing/eval-process/
  5. Anthropic Engineering, "Demystifying evals for AI agents": https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
  6. Simon Willison, notes on "Creating a LLM-as-a-Judge that drives business results": https://simonwillison.net/2024/Oct/30/llm-as-a-judge/
  7. AllthingsPM, "State of AI PM hiring 2026" (604 PM job postings): https://allthingspm.app/blog/state-of-ai-pm-hiring-2026
  8. AllthingsPM AI PM course, Evals chapter: https://allthingspm.app/course/evals
  9. AllthingsPM pricing: https://allthingspm.app/pricing
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free