TAGGED
2 articles

LLM-as-a-judge means using one language model to grade another model's outputs against a written criterion. This guide explains how PMs design, validate and ship a judge, and AllthingsPM teaches it in a graded Evals chapter with interview practice.

AI evals are the tests that decide whether an AI feature is good enough to ship. This guide gives PMs the full method, and AllthingsPM lets you learn it in a graded course chapter and practise it in mock interviews built from real evals job descriptions.
Reading is the easy half.
The course grades the other half.