All Things PM
Daniel Litt: The Mathematician's Guide to AI
The a16z ShowAI

Daniel Litt: The Mathematician's Guide to AI

A working mathematician on where frontier AI already rivals experts and where it still fails, and what happens to a knowledge profession when producing an answer becomes cheap but understanding it does not.

September 1, 2026 · 64 min listen · 11 min read · Daniel Litt
0:00
–:––

Context

Daniel Litt is a mathematician at the University of Toronto who has been unusually vocal about his shifting views on what AI can and cannot do in math. He joins a16z infrastructure partner Leisha Lee to separate the headlines about AI solving hard math from what the models actually do today. The throughline: AI can increasingly produce proofs that would challenge professional mathematicians, but producing a proof is not the same as understanding the problem, and understanding is the thing the work is actually for. Math matters here beyond math itself, because it is one of the first knowledge professions to have a loud, public collision with genuinely capable models, which makes it a preview for every other one, PM work included. These notes pull out the transferable lessons about working alongside AI and about what breaks when output gets cheap, rather than the mathematics.

The Big Idea

As AI makes producing an answer cheap, the scarce and valuable thing shifts to the human understanding and judgment behind it. Any system that still rewards raw output volume will get gamed and stop measuring anything real.

The signal is already visible in math: postdocs can run a model until it emits a correct proof of some old conjecture, so the same theorem shows up proved five different times within days, and low-value papers flood the preprint servers. The output is real; the understanding behind it is not.

Key Insights

1. Solving is not understanding

The goal of the work is not the artifact (the proof, the doc, the code), it is the understanding that produces and survives it. AI can generate a correct result that contains no insight at all.

  • The example: a lemma Litt cared about could be force-proved by a top model, but the output was "10 pages of brutal calculation with no insight whatsoever." Correct, and useless as understanding.
  • Why a PM should care: shipping an AI-generated spec, analysis, or feature is not the same as the team actually understanding the user or the problem. The output can be right while the understanding that lets you build the next thing is missing.

2. Where AI is strong and weak

Litt offers a concrete taxonomy from months of hands-on use, and it generalizes well beyond math.

  • Strong at: applying essentially all known techniques, grinding long technical calculations reliably, pulling together ideas from many papers or areas at once, and running massively parallel work (asking for a thousand examples, or ten sub-agents working ten cases at a time).
  • Weak at: intuition, big-picture "philosophy," building a new theory, deciding which question is even worth asking, and checking an argument's overall structure rather than its line-by-line steps.
  • The pattern: models are excellent at the "last mile" once a problem is well-posed, and weak at the fuzzy, taste-driven front end of figuring out what to work on.

3. Informal reasoning likely generalizes

A common story is that AI is good at math only because math is verifiable, so progress there says little about messier domains. Litt pushes back.

The labs primarily scaled reasoning in natural language, not formally verified (Lean) proofs. Because the thing being scaled is informal reasoning rather than machine-checked logic, his guess is that the same techniques will generalize to other domains that lack a cheap verifier. If true, capability gains in math are a leading indicator for knowledge work broadly, not a special case.

4. Cheap output breaks volume metrics

When producing an output gets cheap, any incentive that counted output stops measuring value, and gets gamed.

  • The disruption shape: a new technology does something slightly worse than before but far cheaper, so a flood of low-quality outputs displaces the previous high-quality ones. Preprint servers have seen exactly this uptick.
  • The gaming: academic incentives reward publishing many papers, so a postdoc can "play the slot machine," running a model until it produces a correct proof of a known conjecture. Litt reproduced this in about an hour: three correct but "quite bad" papers.
  • The tell: the same theorem appearing proved five times within days, because the model mode-collapses onto the same path. Correct, and no evidence a human engaged with any of it.

5. Diversity of exploration is the engine

Progress does not come from one optimal path. It comes from many people chasing their own curiosity, expanding a high-dimensional frontier unevenly, until an idea introduced in one corner suddenly cascades into problems nobody could touch before.

The risk with AI is homogenization. If exploration gets subordinated to whatever the models tend to pursue, you get "one mathematician duplicated a thousand times" instead of a million genuinely different ones. For anyone managing a research or product portfolio, model-driven convergence is a strategic risk, not just an aesthetic one: it quietly narrows the search.

6. Friction can produce the insight

Litt's most useful AI result came from friction, not automation. He had a lemma the models could not prove. Because he could not bring himself to grind out the ugly brute-force version either, he worked examples by hand, found a cleaner restatement, and that better statement had a genuinely beautiful proof (which the model then finished quickly).

His generalization: "our inability to just grind is kind of important to our ability to make discoveries." Removing every bit of friction can also remove the forcing function that produces real understanding. Sometimes the struggle is where the insight comes from, not an inefficiency to engineer away.

7. Verification is the real bottleneck

Models tend to produce short, clever proofs, and it is tempting to read that as taste. Litt's read is bleaker and more useful: they produce short proofs because short is what can be checked.

  • Long arguments are where models fail, because neither the model nor a human can reliably verify them. An 800-page AI-generated proof of a major result is, in his words, almost certainly wrong and unread by anyone.
  • Harnesses that force a model to emit a long result often decrease reliability, because they optimize for producing output rather than for being correct.
  • The lesson for AI product design: the constraint on trusting AI output is verification, not generation. Where you cannot cheaply check the work, long autonomous output is a liability.

Mental Models & Frameworks

Precise question, useful AI

A working heuristic for when to reach for a model: usefulness tracks how precisely the question is posed. The vaguer the phenomenon, the less AI helps.

  • For Litt's deep, multi-year problems, where the hard part is figuring out what the right question even is, AI is "primarily a substitute for Google," useful for learning adjacent topics but not for the real work.
  • For a well-posed lemma where he had already worked out examples, the model proved it fast.
  • Use it as a triage rule: sharpen the question first; the payoff from AI rises steeply once the ask is concrete.

Ascend, don't relinquish

The choice with any capable tool is whether to use it to deepen your own skill or to hand your thinking over to it. Litt describes a "bimodal distribution" in classrooms: some students use AI to level up, others let it do the homework and learn nothing. The same fork faces professionals. The discipline is to use AI to understand more, not to stop understanding, because skill atrophy is the easy default and it is where you lose.

Judge results only in retrospect

You often cannot evaluate how significant a result is until after the fact. Sometimes a problem that looked deep turns out easy; sometimes the reverse. This applies to human and AI output alike, and it is a caution against over-reading any single impressive demo: significance is visible mostly in hindsight, once you see what the idea unlocks.

Trade-offs & Nuance

Optimal is not what happens

Litt grants that even fully superhuman AI might, in principle, be the optimal way to do research. His point is that "optimal" and "what actually happens" are different things. Handing control to models and letting them pursue the direct path gives no guarantee of the broad, diverse exploration that historically drives progress. Keeping capable, curious humans in the loop is the reliable way to steer toward outcomes humans actually want, not an assumption that machines will choose them on their own.

Cheaper and worse, at the same time

The same tool that unlocks huge new reach can lower average quality. AI let Litt take on coding projects he would have procrastinated on for months, a real gain. It also floods his field with correct-but-empty papers. Both are true at once. The resolution is not to reject the tool but to redesign the institutions and incentives around it so the cheap capability is pointed at higher quality, not just higher volume.

Common Mistakes

Playing the slot machine

Running a model repeatedly against a checkable target until it produces something that passes, then claiming credit, without any human understanding of the result. It games output-based rewards, buries real signal under low-value noise, and builds no capability. The fix is to reward engagement and understanding, not artifact count.

Automating away your own understanding

Outsourcing the part of the work that keeps you in the loop. "You don't want to automate your job away, because that does involve you being in the loop to understand it." Convenient in the moment, corrosive over time: you lose the expertise that let you judge the output in the first place.

Practical Application

Use AI on the last mile

Point models at well-scoped, verifiable sub-tasks (a specific lemma, a bounded coding job, a thousand parallel examples) rather than the fuzzy front end of deciding what to pursue. Do your own framing and taste work first, then hand the crisp piece to the model. This matches where it is strong and avoids where it is weak.

Match the tool to question precision

Before reaching for AI, ask how precise your question is. If you cannot state it cleanly, spend the effort sharpening it first, or expect the model to act as little more than a faster search engine. Save the heavy AI lift for the moment the ask is concrete.

Redesign metrics away from volume

Assume any metric that counts output (tickets closed, docs written, features shipped) will be gamed once AI makes that output cheap. Shift measurement toward understanding and outcomes: did the customer problem actually get solved, can the team explain why, did it hold up. Treat volume counts as inputs, not goals.

Protect exploration diversity

If your team leans on the same model for everything, watch for convergence: everyone reaching similar answers because they drew on the same tool. Deliberately preserve independent lines of inquiry and dissenting approaches, the way progress historically depended on many people chasing different intuitions, rather than letting one model's default path quietly narrow the search.

Questions to Consider

  • Where on our team are we rewarding output volume (features shipped, docs written, tickets closed) in a way that AI now makes cheap to game, and what would it mean to measure understanding or outcomes instead?
  • Which of our hard problems are genuinely well-posed enough for AI to help, and which are still at the fuzzy "we do not yet know the right question" stage where a model mostly acts as a search engine?
  • Are we using AI to deepen our own understanding of our users and product, or quietly outsourcing the thinking in ways that will erode the expertise we rely on to judge the output?
  • If everyone on the team leans on the same model, where might our ideas be converging on one narrow path, and how would we notice the loss of diverse approaches?
  • For any AI-generated output we plan to trust, can we actually verify it cheaply, and if not, are we treating a long unverifiable result as if it were reliable?

Bottom Line

Math is the first knowledge profession to publicly meet genuinely capable AI, and its lesson is that solving is not understanding: when producing an answer gets cheap, understanding and judgment become the scarce goods, and any incentive still tied to output volume will break. Use AI on well-posed, checkable tasks to deepen your own understanding, not to replace it, and redesign the metrics and institutions around it before the cheap output games them for you.

Notable Quotes

"The goal of mathematics is not to produce mathematics papers, it's to produce some kind of understanding." (Daniel Litt)

"Our inability to just grind is kind of important to our ability to make discoveries." (Daniel Litt)

"The reason they're not producing long complicated proofs is that they cannot. The ability to check correctness is not yet there." (Daniel Litt)