Flash sale 30% off with code LAUNCH30 Ends in --:--:--
All Things PM
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
The a16z ShowAI Evaluation

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

Vals founder Rayan Krishnan reveals that Meta's Llama 4 underperformed on his private benchmarks while topping the public ones, and explains why enterprises now face an ROI crisis as AI token spend starts to rival employee salaries.

September 9, 2026 · 40 min listen · 12 min read · Ben Horowitz, Rayan Krishnan, Jennifer Li
0:00
–:––

Context

a16z's Ben Horowitz and Jennifer Li sit down with Rayan Krishnan, founder and CEO of Vals (formerly VALS AI), an independent AI model evaluation company started in 2024. The conversation tackles a problem that gets harder as models improve: public benchmarks saturate quickly, get gamed, and increasingly diverge from how a model actually performs on real work. Krishnan explains why self-reported lab benchmarks can't be trusted at face value, walks through how Vals tests a new model in the hours before its public release, and describes a growing enterprise crisis where AI token spend is starting to rival employee salaries with no reliable way yet to measure the return. For a PM, this is a direct look at why "the model scored well on a benchmark" is an increasingly unreliable signal, and what a more rigorous alternative looks like in practice.

The Big Idea

As AI capability outpaces the industry's ability to measure it, independent, continuously evolving evaluation is becoming as essential to AI as auditing is to finance, not just to rank models on a leaderboard, but to let enterprises figure out whether the intelligence they're paying for is actually worth the cost.

Krishnan argues this is compounding into an urgent problem, since some companies are now spending more on AI tokens per employee per month than on that employee's salary, with almost no rigorous way to know if that spend is paying off.

Key Insights

Self-reported benchmarks can hide real underperformance

Krishnan's clearest evidence for why independent evaluation matters: when Meta released Llama 4, the model underperformed on Vals's private, held-out benchmarks while showing what looked like strong capability on major public benchmarks, where the questions and scoring rubrics are openly available. The gap exists because public, open benchmarks can be studied and effectively optimized for, while a model's real capability on genuinely novel tasks may not match. This is why Krishnan believes labs, even when acting in good faith, cannot be the sole source of truth on their own models' capabilities.

Fuzzy human-work distinctions have to become explicit for AI

Krishnan points out an underappreciated problem: humans have never agreed on a clean framework for evaluating human ability either (IQ tests, EQ, the Big Five personality model all have known flaws), and evaluating AI models runs into the same fuzziness, compounded by the fact that models are unusually good at learning to "hack" a known benchmark. His answer is not to find a perfect universal test, but to force previously informal, subjective distinctions, like what separates an associate from a partner at a law firm, into explicit, testable criteria specific to each real workflow. He expects this translation work, making a company's own real work legible enough to test against, to become the biggest long-term bottleneck in evaluation, more than any modeling problem.

Auditor conflicts of interest are the cautionary lesson to avoid

Krishnan draws a direct parallel to the accounting industry's Enron-era failure: when the same firm both audits a company and sells it consulting services, the incentive structure quietly shifts toward "pay to pass the audit." He says other AI benchmarking companies have built a similar conflict by selling training data to the same labs whose models they benchmark, which incentivizes building "gimmick-style" benchmarks as a data-sales channel. Vals made an early, deliberate decision never to sell training data to labs specifically to avoid inheriting this same structural conflict.

Evaluation is shifting from millions of easy samples to few, complex ones

Krishnan describes a structural shift in how benchmarks are built. Early evaluation, exemplified by ImageNet, tested a huge sample size (millions of images) against a simple, one-to-one input-to-label mapping. Modern agentic evaluation flips that ratio: a benchmark might ask a model to build only 50 full-stack web applications, but judging each one now requires a much larger, more complex set of criteria and rubrics. As models take on longer, more open-ended agentic work, this shift toward fewer but richer evaluation samples is likely to keep accelerating, and it also means evaluation infrastructure has to support tasks running over hours, days, or even weeks, including the ability to retry a single failed step without rerunning an entire multi-day trajectory.

Token spend is starting to rival employee salaries

Krishnan shares a concrete anecdote: a Fortune 10 company gave engineers a roughly $100-per-day budget for Claude Code, later increased to $300 per employee, an amount approaching what it might otherwise pay in salary allocation. Because the daily budget reset at a fixed time, employees adapted by concentrating their most productive coding hours right after the reset (4 to 6 p.m. in this example) and taking breaks during the rest of the day when they were rationing usage. Krishnan calls this a sign of "misvaluing of intelligence happening at every layer of the stack," since neither the individual engineers nor the company had a rigorous basis for the budget number, and Anthropic itself is running on narrow margins to serve the underlying compute. His own company hit an extreme version of this directly: during an internal experiment with unlimited coding-tool access, Vals's engineers spent roughly $1.5 million in tokens in one month, about ten times what they spent on salaries for that same period, before the company built internal tooling to rationalize usage.

Recursive self-improvement lacks a shared measurement language

Krishnan describes recursive self-improvement (a frontier model helping train or improve the next version of itself) as one of the most important emerging capabilities to measure, and one of the hardest, since directly having a frontier model train its successor is prohibitively expensive and slow to run as a benchmark. Vals's approach is to build proxy measurements for each discrete part of that pipeline (pre-training, post-training, harness-level engineering) and assess where models show genuine research capability versus where they still struggle, producing what he calls an apples-to-apples way to compare labs on this dimension even though no lab has a standardized way to report it today.

Mental Models & Frameworks

The "always a higher peak" benchmark lifecycle

Vals operates on what Krishnan calls an unofficial motto: always a higher peak. As foundation labs' models saturate a given benchmark (effectively summit that mountain), Vals treats it as its job to perpetually construct the next, harder mountain rather than let labs keep reporting scores against an already-solved test. He extends this with a second, related principle: benchmarks also need periodic recertification against the current state of the world, the same way a lawyer retakes the bar or a doctor renews certification, since a legal research benchmark built on old case law, for example, stops being a meaningful test of current capability even if a model still scores well on it.

Two countervailing forces in AI policy

Krishnan frames AI policy as a tension between two forces that are genuinely difficult to reconcile: moving fast enough that government processes don't slow down beneficial technological progress, and moving carefully enough that the technology's development stays aligned with the public interest. He does not claim a clean resolution, but argues that rigorous, empirical, third-party evidence-gathering (what Vals does) is the mechanism that lets policymakers make that tradeoff deliberately instead of guessing.

Government sets rules, private evaluators verify them

Horowitz offers a specific division-of-labor model for AI policy: the government is well-suited to decide what outcomes it wants to prevent (a model helping with biohacking or cyberattacks, for instance) and to enforce consequences, but poorly suited to the ongoing technical work of actually testing whether a specific model is capable of a specific harmful behavior. He argues that work belongs with competent, independent private evaluators who report findings back to regulators, similar to how a nuclear arms treaty separates the rule (an agreed limit on warheads) from the verification mechanism (an independent inspection, like a flyover).

Trade-offs & Nuance

Independent evaluation slows nobody down, but only if it's fast enough

Krishnan is explicit that Vals's stated goal is to never be "a lagging indicator or a delay to a model release," which means running full evaluation suites across tens of billions of tokens within roughly a six-hour pre-release window. This constraint shapes the whole business: the company had to build massively distributed evaluation infrastructure and an internal automation system (nicknamed Steve) specifically so independent testing doesn't become a bottleneck labs would be tempted to skip under launch-timeline pressure.

Sovereign AI development undercuts a unified evaluation standard

Krishnan admits that, from an idealistic efficiency standpoint, the amount of duplicated investment in country-specific ("sovereign") AI, separate data centers, separate large models, separate training pipelines, looks wasteful compared to a more consolidated global effort. But he acknowledges that is not the world being built, and different countries embed different values directly into their models (he cites Chinese open source models being restricted on discussing certain historical and political topics). His conclusion is that a shared evaluation language becomes even more necessary, not less, specifically because sovereign AI development is fragmenting rather than converging.

Practical Application

Build an internal benchmark from your own codebase before picking a coding model

Krishnan describes Vals's own product, ValSmith, built directly out of the company's internal token-spending crisis: it lets a company take its actual GitHub codebase and construct a benchmark specific to that codebase, to identify which coding agent is actually the highest-ROI choice for that repository rather than assuming the best-known public model is automatically the best fit. Vals found real, non-obvious results this way, including that Anthropic's Sonnet model was sometimes more expensive in practice than the ostensibly higher-tier Opus model, because Sonnet consumed tokens at a much higher rate for comparable tasks.

Route routine and high-difficulty tasks to different tools deliberately

Rather than giving every team unrestricted access to every AI tool and hoping usage self-regulates, Vals moved to auto-issuing a recommended starting tool for each GitHub issue or ticket, calibrated to the task's actual intelligence requirements, while still leaving engineers flexibility to escalate to a more expensive tool for genuinely harder work. This followed directly from discovering that one engineer alone spent six billion tokens in a single peak day during an unrestricted-access experiment.

Treat a benchmark's provenance as part of your evaluation criteria

Before trusting a public benchmark score when choosing a model or vendor, check who built and maintains that benchmark and whether they have a financial relationship with the entity being scored, such as selling training data back to the labs. Krishnan's account of "gimmick-style" benchmarks built specifically as a sales channel for training data is a concrete warning sign to watch for when a vendor cites benchmark performance as its main proof point.

Expect agentic evaluation to require session-level infrastructure, not one-shot scoring

If you are building or buying evaluation tooling for agentic AI systems that run over long horizons, make sure the infrastructure can retry a single failed step in a multi-hour or multi-day task without restarting the whole trajectory, and that scoring criteria account for a small number of complex tasks rather than assuming a large sample of simple, one-shot questions.

Questions to Consider

  • If your organization has adopted an AI coding tool based on a public benchmark score, have you validated that score against your own actual codebase and workflows, the way Vals's ValSmith product lets a company build a benchmark from its own GitHub repository?
  • Does your company have a clear, evidence-based basis for its per-employee AI token budget, or is it, like the Fortune 10 example Rayan Krishnan describes, closer to an arbitrary number that has already needed to be raised once?
  • If a vendor providing your AI evaluation or benchmarking data also sells services or training data to the labs it evaluates, does that create the same kind of conflict of interest that undermined pre-Enron accounting audits?
  • As agentic AI systems take on tasks that unfold over days or weeks, does your team's current way of measuring "is this working" still make sense, or was it designed for single-shot question-and-answer interactions?

Bottom Line

Rayan Krishnan's core argument is that public, self-reported AI benchmarks are becoming actively misleading as models get better at optimizing for known tests, which makes independent, continuously evolving evaluation essential both for labs proving real progress and for enterprises trying to justify AI spend that increasingly rivals payroll. The company or team that makes its own real-world work legible enough to evaluate against, rather than relying on generic leaderboards, will be the one that actually knows whether its AI spend is paying off.

Case Studies Mentioned

Meta's Llama 4 benchmark disconnect

When Meta released Llama 4, Vals's private, held-out benchmarks showed the model underperforming, while the same model appeared to show strong capability on major public benchmarks, where the underlying questions and scoring rubrics are openly available and can be studied in advance. Krishnan cites this as the clearest public example of why self-reported or openly-gameable benchmark scores can diverge sharply from a model's real capability, and why he believes the industry needs evaluators with no stake in how well a given model scores.

Vals's own $1.5 million token-spending month

During an internal experiment giving the Vals team unlimited access to coding tools for a month, engineers collectively spent roughly $1.5 million worth of tokens, about ten times the company's salary spend for that same period, with one engineer alone using six billion tokens on a peak day. Rather than continuing with unrestricted access, the company built its own internal tool, ValSmith, to analyze which coding agents were actually token-efficient for its workflows (finding, for example, that the Cognition Devin tool was notably efficient) and used that data to set smarter usage guidance going forward.

People to Follow

Rayan Krishnan

Founder and CEO of Vals, an independent AI model evaluation company started in 2024 after his research background in building benchmarks convinced him public benchmarks could no longer reliably measure model progress. He has since built out Vals's benchmark catalog (finance, coding, legal research, and a recursive self-improvement index among them) and regularly briefs government executive and legislative branches on AI capability and risk findings.

Ben Horowitz

Co-founder of Andreessen Horowitz, who brings a policy-and-history lens to the conversation, drawing comparisons between AI evaluation's current ambiguity and historical analogs like the MPAA's fuzzy content-rating standards and the accounting industry's conflict-of-interest failures around the Enron collapse. He also spends significant time engaging directly with policymakers in Washington, D.C.

Notable Quotes

"Every time a new trillion dollar industry emerges, there's a need for this independent testing group." (Rayan Krishnan)

"There is a misvaluing of intelligence happening at every layer of the stack." (Rayan Krishnan)

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free