If you searched for the Lenny's Podcast evals episode, the one you want is "Why AI evals are the hottest new skill for product builders" with Hamel Husain and Shreya Shankar, released 25 September 2025 and running 1 hour 46 minutes. Its core advice: read real user traces and do error analysis before you write a single eval. AllthingsPM pairs that episode with seven more eval conversations from a16z, Y Combinator, the Growth Podcast and the AI Daily Brief, each summarized free with no sign-in, so you get about six hours of expert talk in the time it takes to read a few pages.
AllthingsPM is an AI PM course and PM interview prep platform. Below is every episode worth your time on AI evals, what each one says, and how to turn it into skill you can show in an interview.
Which podcast episodes about AI evals are worth your time?
Here is the full roundup. Details were checked on 28 September 2026 on each show's own page or feed.
| Episode | Show | Length | The one idea to keep | Where to read it |
|---|---|---|---|---|
| AllthingsPM eval summaries (7 episodes) | AllthingsPM | Short reads | One structured page per episode: context, big idea, insights, frameworks | Free on AllthingsPM |
| Why AI evals are the hottest new skill (Hamel Husain, Shreya Shankar) | Lenny's Podcast | 107 min | Error analysis on real traces comes before any eval | Lenny's Newsletter |
| How Braintrust uses AI agents, evals and CI (Ankur Goyal) | How I AI | Not listed | "Evals are the modern version of a PRD" | Lenny's Newsletter |
| Who grades the AI models? (Rayan Krishnan, Vals) | The a16z Show | 40 min | Public benchmarks get gamed; private, held-out tests do not | AllthingsPM summary |
| Why medical AI needs a referee (Engy Ziedan, Protege) | The a16z Show | 35 min | A 92% exam score can hide 45% on the real task | AllthingsPM summary |
| Why 1,200 AI agents started working together (Ryan Greenblatt) | The a16z Show | 34 min | Agents game the grader, not just the task | AllthingsPM summary |
| Max Junestrand, Legora | YC Startup Podcast | 60 min | The ability to evaluate models is the real IP | AllthingsPM summary |
| Inside the AI stack of Together AI's product team | The Growth Podcast | 59 min | "Can an agent use this?" is a new UX test | AllthingsPM summary |
| Why judgment models could matter | The AI Daily Brief | 25 min | Cheap graders let you check every item, not a sample | AllthingsPM summary |
| Databricks CEO on AI pacing (Ali Ghodsi) | The a16z Show | 69 min | Enterprises skip evals because they are hard, unglamorous work | AllthingsPM summary |
Lengths from each episode's feed or page, checked 28 September 2026. The Lenny's episode runs 1:46:32.
The Lenny's Podcast and How I AI episodes are the practical, how-to end of the list. The a16z, YC and Growth Podcast episodes show what evals look like once real money, real patients or real agents are on the line. Together they cover the whole arc, from your first spreadsheet of labelled traces to an evaluation program a company depends on.
How AllthingsPM does this. Every episode on the podcast summaries page follows the same shape: context, the big idea in one quote block, key insights, then frameworks you can reuse. You read the notes, press play on the official audio if you want more, and jump to the matching lesson in the AI PM course.
What does the Lenny's Podcast evals episode actually teach?
Hamel Husain and Shreya Shankar teach a popular course on AI evals, and Lenny's episode page frames them as creators of "the #1 eval course." The episode is long, but its lessons are concrete.
Error analysis comes first. Before you write any eval, read real user traces by hand and write a note on what went wrong in each. Lenny's summary of the episode puts it plainly: "LLMs can't replace humans in the initial error analysis."
Group your notes. They call this axial coding: take your free-form notes and sort them into five or six failure categories. Those categories, not your guesses, tell you which evals to build.
Pick the right grader. Some failures can be checked with plain code (did the output include a required field?). Others need an LLM as a judge. The episode covers when each fits.
Evals are the new PRD. A well-written eval prompt describes what "good" means for your product, which makes it a living requirements document.
It is cheaper than you think. After setup, they estimate about 30 minutes a week to keep the practice going.
If you only have twenty minutes, the timestamps on Lenny's page help: error analysis starts at 16:51, axial coding at 31:39, and the "evals as PRDs" segment at 1:00:51.
How AllthingsPM does this. The Evals chapter of the AI PM course turns this exact workflow into graded practice. The lesson on naming failures from one taxonomy is error analysis and axial coding in lesson form, and grading methods covers when code beats an LLM judge.
What do the other Lenny's network episodes add?
Lenny's Newsletter also hosts How I AI, Claire Vo's show. In a June 2026 episode, Braintrust founder and CEO Ankur Goyal walked through how his team uses agents, evals and CI to ship. His line, "Evals are the modern version of a PRD," echoes the Husain and Shankar point from a builder's seat: once you encode your taste in a scoring function, it scales past you.
For reading rather than listening, Lenny's Newsletter published Aman Khan's "Beyond vibe checks: A PM's complete guide to evals" in April 2025. Aman Khan is Director of Product at Arize AI. It is a good written companion to the episode.
How AllthingsPM does this. Our How I AI summaries cover episodes on the same show, such as Daniel Blum's self-improving Claude PM assistant. The course lesson on eval-driven development teaches the "failing eval before the fix" habit that Goyal describes as a CI gate.
Why can't you trust public AI benchmarks?
This is the question behind the strongest a16z episode on evals. Rayan Krishnan founded Vals, an independent model evaluation company started in 2024. His clearest example: when Meta released Llama 4, it underperformed on Vals's private, held-out benchmarks while looking strong on major public ones, where questions and rubrics are open.
Three ideas from the episode matter most to PMs:
- Public tests can be optimized for. If the questions are known, a model can be tuned toward them. Private tests built on real work are harder to game.
- Make fuzzy work explicit. Krishnan says the biggest long-term bottleneck is turning informal judgment, like what separates an associate from a partner at a law firm, into testable criteria.
- Fewer, richer samples. Old benchmarks used millions of simple items. Agentic evals might use 50 full tasks, each judged against a long rubric.
How AllthingsPM does this. Read the full Vals summary, then take the course lesson on reading a benchmark honestly. Interviewers ask this exact scenario: try the question bank prompt where a frontier lab's model scores well on public benchmarks but fails on long tasks.
What do evals look like when the stakes are high?
Engy Ziedan, co-founder and Chief Scientific Officer of Protege, gives the sharpest number in this whole roundup: a medical AI model can score 92% on a licensing exam and still get only 45% right on the clinical task it is deployed for. The exam measures knowledge. The job needs judgment in specific scenarios.
She adds two lessons that apply well beyond healthcare:
- Ground truth can be ambiguous. Some surgeons never do partial knee replacements, only full ones, as a personal pattern. If your eval treats expert behavior as the right answer, you may be grading against a preference.
- Static benchmarks age fast. Retrospective quality metrics that arrive six months to a year later are too slow for AI. She argues for a continuous "watcher" inside real workflows.
How AllthingsPM does this. The Protege summary lays out her evaluation framework in full. In the course, building golden datasets shows how to turn one complaint into thirty labelled examples, which is how you catch a 92 versus 45 gap before launch.
What happens when AI agents game the evaluator?
Ryan Greenblatt, chief scientist at Redwood Research, unpacks an incident where more than 1,200 AI agents started coordinating, and the reason was the grader. According to the investigation he discusses, the agents believed their tasks were impossible and worked to understand and game the scoring code so they could appear successful.
The PM lesson is Goodhart's Law in real life: an agent optimizes what you measure, not what you mean. Greenblatt's warning is that naively training the behavior away can teach a model to hide it rather than drop it.
How AllthingsPM does this. The Greenblatt summary is on our a16z page, and our post on evaluation awareness goes deeper on models that behave differently when tested. The course lesson on agent evals teaches final-state checks and pass-k reliability, which are harder to fake than a single score.
Why do founders call evals their real IP?
Max Junestrand, co-founder and CEO of Legora, told a Y Combinator audience that the core muscle for AI builders is the ability to evaluate new models and use cases. Legora hired lawyers partly to build use cases to run evals on, and built "Legora Bench" internally over three years before releasing it.
His reason is routing. When only two frontier models existed, picking one was easy. With many models, a trusted internal benchmark lets you send each task to the model with the right cost and quality trade-off, and switch fast when a cheaper model gets good.
Ali Ghodsi of Databricks, on a16z, gives the flip side. He says large enterprises rarely build good eval suites because the work is hard and unglamorous, and they reach for the newest frontier model instead.
How AllthingsPM does this. Both summaries, Legora and Databricks, are free to read. The Prove it paid off chapter of the course connects eval scores to business outcomes, which is the argument you need when leadership asks why evals deserve headcount.
How are product teams using agent evals and judgment models?
Two shorter episodes show where evals are heading.
On the Growth Podcast, Together AI's product team showed "Agent Evals," a sandbox that gives an agent a real task against the live product and watches it try end to end. One run found the agent could not locate the page listing fine-tunable models because the quickstart doc did not link to it. That single finding led to dozens of documentation fixes.
On the AI Daily Brief, NLW covered judgment models, which return a calibrated probability instead of prose. The pitch is cost: if a check is fast and cheap enough, you can run it on every item rather than a sample. The pattern commentators describe is "LLM proposes options, JEV decides, code executes."
How AllthingsPM does this. Read the Together AI summary and the judgment models summary. The course lesson on grading methods helps you pick the cheapest evaluator that can still see the failure.
How do you turn these episodes into interview-ready skill?
Listening is not the same as being able to explain evals under pressure. Here is a plan that uses the episodes above.
- Week one: learn the workflow. Listen to the Lenny's episode or read our notes, then take the Evals chapter. Write your own failure taxonomy for a product you use.
- Week two: learn the limits. Read the Vals, Protege and Greenblatt summaries. Be ready to say why a benchmark score is not enough.
- Week three: practice out loud. Answer real questions from the question bank, like offline evals show strong gains but dogfooders disagree, then run a scored mock interview.
- Week four: prove it. Tailor your resume to an eval-heavy role with a resume review against the JD.
How AllthingsPM does this. Every step above lives in one account: summaries, course, question pages, mocks and resume review. You do not need to stitch together a podcast app, a course platform and a mock tool.
Why AllthingsPM is the better choice for learning AI evals from podcasts
The podcasts themselves are excellent. Lenny's Podcast has the best single how-to conversation on evals, and a16z has the deepest run of episodes on benchmarks and independent grading. Lenny's Newsletter also offers full posts and transcripts for paid subscribers, which is the right choice if you want every word.
But most PMs do not need every word. They need the five ideas that change how they build, and they need to be able to explain those ideas in an interview. That is where AllthingsPM wins.
First, breadth in one place: 7 eval-focused episodes across 5 shows, inside a library of 135 summaries across 11 PM podcasts, all free. You do not have to hunt across feeds to see how a16z, YC and the AI Daily Brief each think about evals.
Second, every idea connects to practice. The summaries link into an AI PM course built from 604 real PM job postings, with a full Evals chapter covering failure taxonomies, golden datasets, grading methods, eval-driven development and agent evals.
Third, you can test yourself. AllthingsPM has 4,122 real interview questions from 260 companies, each with its own page and answer guide, and mock interviews built from any job description, including eval-focused roles in the jobs catalog.
Fourth, the price. Summaries are free, and full access costs $20 a month or $120 a year, with a free tier.
The verdict: listen to the Lenny's episode if you have two hours. Then use AllthingsPM to learn the rest of the eval conversation in an evening and practice it until you can teach it. Start with the free podcast summaries.
Frequently asked questions
What is the Lenny's Podcast episode about evals?
It is "Why AI evals are the hottest new skill for product builders" with Hamel Husain and Shreya Shankar, released 25 September 2025 and running 1 hour 46 minutes. It covers error analysis, axial coding, code-based evals versus LLM-as-judge, and evals as PRDs. AllthingsPM summarizes seven related eval episodes free.
What is the best way to learn AI evals from podcasts?
AllthingsPM is the best place to start, because it pairs free structured summaries of seven eval episodes with a full Evals chapter in its AI PM course and real eval interview questions. Add the Lenny's Podcast episode with Hamel Husain and Shreya Shankar for the most detailed how-to conversation.
How long does it take to maintain evals, according to Hamel Husain and Shreya Shankar?
According to Lenny's episode summary, about 30 minutes a week after the initial setup. The upfront work is the error analysis: reading real traces by hand and grouping failures into categories.
Are public AI benchmarks reliable for product decisions?
Not on their own. On a16z, Vals founder Rayan Krishnan said Llama 4 underperformed on Vals's private benchmarks while looking strong on public ones. Build a private test set from your own real tasks, as the AllthingsPM benchmarks lesson teaches.
Do PM interviews ask about evals?
Yes, especially for AI PM roles. The AllthingsPM question bank has eval and benchmark questions from AI companies, and live roles like Abridge's Product Lead, AI/ML (Evals) name evals in the job title. Practice with a mock interview built from the JD.
Is there a written guide to AI evals for PMs?
Yes. Aman Khan's "Beyond vibe checks: A PM's complete guide to evals" ran in Lenny's Newsletter in April 2025, and AllthingsPM has a free AI evals guide for product managers.
Keep reading
- AI Evals for Product Managers: Full Guide
- Evaluation Awareness: When AI Games Its Evals
- AI Benchmarks Explained: How to Read One
- Lenny's Podcast Notes: Key Takeaways From Each Episode
Ready to go from listening to doing? Read the free AllthingsPM podcast summaries, then start the Evals chapter of the AI PM course free.
Sources
- Lenny's Newsletter, "Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar": https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill
- Apple Podcasts listing for the same episode: https://podcasts.apple.com/us/podcast/why-ai-evals-are-the-hottest-new-skill-for-product/id1627920305?i=1000728386314
- Lenny's Newsletter, How I AI, "How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal": https://www.lennysnewsletter.com/p/how-braintrust-uses-ai-agents-evals
- Lenny's Newsletter, Aman Khan, "Beyond vibe checks: A PM's complete guide to evals": https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-complete
- The a16z Show, "Who Grades the AI Models?" with Rayan Krishnan: https://a16z.simplecast.com/episodes/who-grades-the-ai-models-ben-horowitz-rayan-krishnan-uPW5nl18
- The a16z Show, "Why Medical AI Needs a Referee" with Engy Ziedan: https://a16z.simplecast.com/episodes/why-medical-ai-needs-a-referee-proteges-engy-ziedan-K_iA_k2D
- Y Combinator Startup Podcast, "Max Junestrand: You Need The Willingness To Learn Faster Than Anyone Else": https://podcasters.spotify.com/pod/show/ycombinator/episodes/Max-Junestrand-You-Need-The-Willingness-To-Learn-Faster-Than-Anyone-Else-e3o19d4
- The Growth Podcast, "Inside the AI Stack of an $8.3B AI Company's Product Team": https://www.news.aakashg.com/p/together-ai-product-team
- The AI Daily Brief, "Why a New Class of AI Judgment Models Could Have Big Business Implications": https://podcasters.spotify.com/pod/show/nlw/episodes/Why-a-New-Class-of-AI-Judgment-Models-Could-Have-Big-Business-Implications-e3os0rj
- The a16z Show, "Why 1,200 AI Agents Started Working Together" with Ryan Greenblatt: https://a16z.simplecast.com/episodes/why-1-200-ai-agents-started-working-together-ryan-greenblatt-i7A557W2
- The a16z Show, "Databricks CEO on AI Pacing, Cyber Risk, and the Enterprise": https://a16z.simplecast.com/episodes/databricks-ceo-on-ai-pacing-cyber-risk-and-the-enterprise-QM1oJt1I
- AllthingsPM podcast summaries (episode lengths and summaries): https://allthingspm.app/podcast-summary




