All Things PM
Why Medical AI Needs a Referee | Protege's Engy Ziedan
The a16z ShowResearch

Why Medical AI Needs a Referee | Protege's Engy Ziedan

Protege's Engy Ziedan explains why a medical AI model can score 92% on a licensing exam and still only get 45% right on the actual clinical task it's deployed for, and why nobody in healthcare AI is positioned to catch that gap except an independent referee.

August 24, 2026 · 35 min listen · 10 min read · Engy Ziedan
0:00
–:––

Context

a16z's Daisy Wolf and Eva Steinman interview Engy Ziedan, co-founder and Chief Scientific Officer of Protege and a healthcare economist, about why medical AI evaluation is fundamentally broken. Protege started as a healthcare data provider (most foundation models are pre- and mid-trained on its data) and has since moved into building independent, continuous evaluations for vertical medical AI tools, arguing that acing a standardized exam doesn't mean a model is ready to make real clinical decisions. For a PM building or buying AI in any high-stakes domain, not just healthcare, the episode is a detailed case study in why benchmark performance and real-world task performance are different things, and what it actually takes to measure the difference honestly.

The Big Idea

A medical AI model can score 92% on a licensing exam and still only get 45% right on the actual clinical task it's deployed for, because standardized benchmarks measure general knowledge while real clinical work requires judgment in specific, high-stakes scenarios that static test questions don't capture, and no one in the industry is currently positioned to measure that gap independently.

Ziedan argues the industry has plenty of effort going into catastrophic-failure safety testing, but almost no one is positioned to catch the harder, subtler problem: models that perform well on paper while being misaligned with the patient's actual interest in a live clinical workflow.

Key Insights

Benchmarks measure knowledge, not job performance

Ziedan draws a direct analogy to hiring: a licensing-exam-style benchmark is like screening a candidate on a written test, while the real question for deployment is whether the model has actually performed well in the specific scenario it will be used for. Her example: a model that passed a 5,000 to 10,000 question medical exam is being compared, functionally, to a physician who has performed thousands of a specific procedure. The number that matters for a patient facing spinal surgery isn't general medical knowledge, it's whether the tool has been used in that exact scenario before and what happened when it was.

Catastrophic failure is easier to catch than subtle misalignment

Ziedan distinguishes two categories of AI safety risk in healthcare: catastrophic failure (a narrowly definable outcome like mortality, which is comparatively easier to test for) and subtle misalignment or bias, which is harder to define and even harder to detect. She illustrates this with a prior-authorization example: insurers deploy agentic AI to minimize payouts while hospitals deploy agentic AI to maximize revenue recovery, both technically "working" for their side, while the patient in the middle has no equivalent agent and no regulator is verifying whether either system is actually serving the patient's interest, even though both models could look high-performing on their own success metrics.

Physician preference makes ground truth genuinely ambiguous

Real-world clinical benchmarking runs into a problem beyond model accuracy: physicians themselves have consistent individual practice preferences (a "beat," in Ziedan's term) that don't necessarily reflect a single correct answer. Her example: some orthopedic surgeons never perform partial knee replacements, only full ones, and vice versa, as a matter of personal practice pattern rather than clinical necessity. When a model's recommendation on a real patient record diverges from what the treating physician actually did, that divergence could mean the model is wrong, or it could mean the physician's own practice pattern was never a clean ground truth to begin with, which complicates any evaluation that treats physician behavior as the correct answer.

AI can reduce documentation bias, not just introduce it

Ziedan offers a counterexample to the assumption that AI mainly introduces bias into clinical care: ambient scribing tools that transcribe patient-physician conversations directly into clinical notes can strip out the kind of subjective, bias-carrying language that historically appeared in manually written notes, phrases like "she looks disheveled" or "her husband is asking good questions," language Ziedan says has historically signaled a note-writer's judgment that a patient isn't trustworthy and led to symptoms being dismissed as psychological. A more literal, transcription-based note can prompt a third-party physician reviewing the same case to order the physical workup the original bias-inflected note skipped. She argues this benefit isn't currently being measured any more than the harms are, evaluation is missing on both sides.

Vendor self-reported evals create a market-wide information asymmetry

Every vertical AI vendor in a given clinical subdomain (ambient scribing, prior auth, oncopathology, nurse workflows) claims to be best in its own marketing material, but Ziedan says there's no incentive for any vendor to publish where its own product actually fails, and no independent party currently positioned to verify competing claims. She frames this using a classic information-asymmetry problem: buyers (hospitals, pharma companies, health systems) are being asked to spend millions of dollars onboarding a tool without a reliable, comparable way to know which vendor's claims are accurate for their specific use case.

Mental Models & Frameworks

Hedonic value and why evals matter economically

Ziedan applies an economic concept from her background as a healthcare economist: hedonic pricing infers the value of a good from what people are willing to pay for it (she cites school districts pricing into housing values as an example), and notes that historically, consumer products like Uber never needed a formal "eval" because people simply paid for rides and the value was self-evident through usage. AI is different, she argues, because the technology doesn't have a marginal cost of zero (people pay in tokens, onboarding time, and the risk of something going wrong), so accurate pricing and market clearing require the kind of institutional trust that only comes from robust, third-party evaluation, the same way you'd vet a new employee before hiring them, not just after they start causing problems.

Continuous monitoring versus retrospective quality reporting

Ziedan contrasts two evaluation cadences using the government's Value-Based Purchasing program as a historical reference point: that program ties a large share of government healthcare payments to quality metrics, but those metrics (like penalizing hospitals for antibiotic overprescription or nursing home fall rates) are measured retrospectively, often six months to a year after the fact. She argues this cadence, while appropriate for slower-moving human care systems, is fundamentally too slow for AI, where models can face data drift, per-physician personalization, or continuous test-time learning, meaning a static benchmark published even a year prior may already be measuring a system that no longer exists. Her proposed alternative is a live "watcher," continuous evaluation embedded in real clinical workflows rather than periodic testing against a fixed question set.

Trade-offs & Nuance

Independent evaluator with a vendor's own data supply

Ziedan directly addresses the apparent conflict of interest in Protege positioning itself as a neutral evaluator while also being a major data supplier to the same foundation model labs it evaluates. Her defense rests on two points: a comparative-advantage argument (Protege's core capability, having the scale and diversity of healthcare data needed to both diagnose a model's specific weakness and supply exactly the data that would fix it, is what makes it uniquely positioned, not conflicted) and a trust-based network-effects argument (since the evaluation business depends entirely on being seen as impartial by every vendor being compared, any breach of that trust would collapse the business immediately, which she argues is itself a strong incentive toward genuine neutrality). The unresolved tension buyers should weigh for themselves: a referee that also sells training data to the teams it referees has a different incentive structure than a fully arm's-length auditor, even if the business logic for why it stays honest is coherent.

Common Mistakes

Mistake: treating a static exam-style benchmark as proof of real-world readiness

Ziedan's core critique is that scoring well on a fixed set of exam-style questions demonstrates general knowledge retention, not competence at a specific real-world task under real-world conditions. The fix isn't abandoning benchmarks, it's recognizing they answer a narrower question ("does this model know things") than the one that actually matters for deployment ("has this model performed well in this specific high-stakes scenario before").

Mistake: assuming evaluation data hasn't already leaked into training

Ziedan describes discovering that roughly 80% of Protege's own healthcare data had already been supplied for model training somewhere, making it much harder than expected to build genuinely uncontaminated evaluation sets. Protege now maintains what she calls a sealed membrane, a strict rule that any patient data used in training is permanently excluded from ever being used in a benchmark or evaluation, sourcing net-new data (like never-before-scanned pathology slides) directly from hospitals specifically to keep evaluation sets clean. The broader lesson: assume contamination is more pervasive than expected, and build an explicit, enforced separation between training data and eval data rather than trusting a benchmark's provenance by default.

Practical Application

Ask any AI vendor what task-specific track record backs their benchmark score

Before trusting a vendor's benchmark claim, particularly in a high-stakes domain, ask the equivalent of Ziedan's spinal-surgery question: not "how did this model do on a general knowledge test," but "how many times has this specific model been used for this specific task, and what happened." A high general benchmark score is a necessary but insufficient signal for deployment readiness in a narrow, high-consequence use case.

Build (or demand) continuous evaluation, not one-time certification

If you're deploying an AI system in a domain where the underlying model, its context, or its training data can drift over time, treat a one-time evaluation report as a snapshot with a shelf life, not a permanent certification. Ziedan's argument for a live "watcher" applies beyond healthcare: any AI system operating in a live workflow where humans can inadvertently shift its behavior (she gives the example of a group of nurses whose comments about a patient nudge a model toward decisions that favor their own working hours over the patient's care) benefits from ongoing monitoring rather than periodic retesting alone.

Questions to Consider

  • If our own AI product's marketing benchmark score is based on general knowledge or exam-style questions, do we actually know how it performs on the specific, narrow task our customers use it for in practice?
  • Are we confident that our evaluation or test data has never been part of our model's training data, or could contamination be quietly inflating scores we're relying on to make deployment decisions?
  • If a group of end users could inadvertently shift our AI system's behavior through how they interact with it, the way Ziedan describes nurses' comments nudging a model's recommendations, would we currently have any way to detect that drift before it caused harm?

Bottom Line

A high score on a general benchmark tells you a model knows things, not that it's safe or effective at the specific, high-stakes task you're deploying it for, and because no vendor is incentivized to publish where its own product fails, real evaluation has to come from an independent party doing continuous, task-specific testing rather than a one-time exam-style score.

Concepts to Explore

Information asymmetry in healthcare markets

Ziedan cites research on the effect of having a doctor in the family (a Swedish study using random assignment via medical school admission cutoffs) showing family members live measurably longer, as evidence of how much more providers know about care value than patients do. This asymmetry, well studied in healthcare economics generally, becomes more acute once AI agents are deployed inside that same information gap on behalf of either providers or payers, since patients have even less visibility into how an algorithmic decision was actually made than into a human one.