Flash sale 30% off with code LAUNCH30 Ends in --:--:--
All Things PM
Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology
All-In with Chamath, Jason, Sacks & FriedbergAI Infrastructure

Naveen Rao: 4D Computing, AI's Energy Wall & Beating Biology

A squirrel's brain runs flawless real-time motor control on 8 nanowatts. Naveen Rao's new chip company built a working prototype in five months by asking why computers move so much more data than brains do.

September 21, 2026 · 23 min listen · 7 min read · Naveen Rao
0:00
–:––

Context

Naveen Rao, founder of Nervana Systems (an early AI chip company acquired by Intel) and MosaicML (acquired by Databricks), presents his new company Unconventional AI at the All-In Summit, arguing that AI's real near-term constraint is energy, not chip count, and that the fix requires rethinking computer architecture itself rather than continuing to scale existing designs. The episode matters to PMs and technical leaders working anywhere near AI infrastructure because Rao makes a specific, quantified argument for why current hardware architecture (not just insufficient chip supply) is the actual bottleneck, and demonstrates a working alternative built in five months.

The Big Idea

The energy cost of running AI models comes overwhelmingly from moving data between memory and compute, not from the computation itself, and Rao's company built a fundamentally different chip architecture, one where memory and compute are the same physical element rather than separate components connected by a data bus, specifically to eliminate that movement cost, producing a claimed 1000x power efficiency improvement over conventional GPUs.

His clearest supporting comparison: a GPU moves roughly 30 trillion bits per second in and out of memory, while the human cortex, which has 13-14 billion neurons, moves only about 16 billion bits per second, a difference of several orders of magnitude that Rao attributes directly to biological systems co-locating memory and computation rather than shuttling data between separate components.

The Core Argument

Energy, not chip supply, is the binding constraint on AI's growth curve

Rao's specific arithmetic: Google alone processes roughly 3.2 quadrillion tokens per month, and even using a conservative 10 joules-per-token estimate, that works out to about 12 gigawatts of continuous power draw for one company's AI services, against roughly 40 gigawatts of total US data center capacity and under 100 gigawatts worldwide. His extrapolation: if model size and demand both keep growing (which they are), available energy runs out on a roughly three-year horizon, well before chip manufacturing capacity would become the limiting factor. This reframes a common industry assumption (that GPU supply or fab capacity is the primary bottleneck) by putting a specific number on a different, earlier-arriving constraint.

Roughly half the cost of serving a single AI query is now energy, not hardware

Rao states that approximately 50% of the cost of serving a token (a single unit of AI output, like one response to a ChatGPT query) is energy, with the remainder split between hardware capital expenditure and facility costs. His broader observation about how data center planning has changed: the sequence used to be "get floor space, then networking, then GPUs," and today it starts with "get an energy contract, then figure out how to build infrastructure to monetize every watt of it," a genuine reordering of the constraint hierarchy that changes what an infrastructure buyer should actually be negotiating for first.

Biological brains demonstrate the efficiency ceiling is roughly ten billion times higher than current hardware

Rao's specific biological benchmarks: the human brain runs on about 20 watts; a scaled-down animal brain (his example, a monkey) runs on roughly 1 watt, comparable to a smartphone; and a squirrel's brain, capable of precise, repeated, high-stakes motor tasks (accurately jumping between branches essentially every time), runs on about 8 nanowatts, meaning you could run more than 100 squirrel-brain-equivalents on the power budget of a single phone. His claim: current AI hardware sits roughly ten orders of magnitude away from the thermodynamic efficiency ceiling that biological neural systems already operate within one or two orders of magnitude of, framing the gap not as a modest engineering improvement opportunity but as a multi-order-of-magnitude one.

The fix requires removing the memory-compute separation that's defined computers since the 1940s

Rao traces a consistent architectural pattern from the 1945 ENIAC through modern GPUs: a separate memory store and a separate compute unit, with data continuously shuttled between them, an approach optimized historically for raw speed (ENIAC was built to be faster than human artillery-trajectory calculators) without ever optimizing for energy efficiency, because energy was never the binding constraint until now. His company's core architectural bet is a "dynamical computer" where each physical computing element is itself a memory element, eliminating the separate memory interface entirely, which he calls "4D computing" (three physical dimensions of chip stacking plus a time dimension in how the system's physical dynamics evolve) as a genuinely different category from the CPU-to-GPU evolution, which he notes is still fundamentally the same "von Neumann architecture" underneath.

A working prototype and a public model demonstrated the concept before claiming the full efficiency target

Rao describes a specific execution timeline: the company started in earnest in January, taped out (finalized and sent to a chip fabrication facility) its first physical chip design by June 1st of the same year, and had working results back from the fabricated chip generating images at roughly 500 nanojoules per image, compared to the millijoule range typical of a conventional GPU, several orders of magnitude more efficient on that specific task. Separately, the company released an open-source generative image model called UNO built on simulated oscillator-based dynamics (physically inspired by how multiple metronomes on a shared, free-moving platform will spontaneously synchronize through pure physical coupling) specifically to demonstrate the underlying computational principle before the physical chip existed, a sequencing choice worth noting: prove the computational concept in simulation first, then build the physical hardware to validate the efficiency claim.

Sparsity improved both efficiency and trainability simultaneously, a genuinely rare combination

Rao describes discovering that deliberately removing connections between computing elements (sparsity), which avoids the quadratic cost growth of connecting every element to every other element, didn't just reduce cost, it also made the resulting system easier to train and produced better performance. He frames this explicitly as an unusual outcome: normally efficiency gains and performance are in tension, but this particular architectural choice improved both simultaneously, which he credits to reframing the underlying problem rather than optimizing within the existing framing.

Practical Application

Separate the "is this energy-bound or compute-bound" question before planning any AI infrastructure investment

Before committing capital to AI infrastructure, explicitly check whether your actual constraint over the relevant time horizon is compute/chip availability or power availability, following Rao's argument that energy, not chip supply, is likely to bind first. The right infrastructure strategy differs substantially depending on which constraint you're actually planning against.

Look to biological or physical systems for an efficiency benchmark before assuming current engineering approaches are near their ceiling

When evaluating whether an engineering approach in your own domain is close to its practical efficiency limit, check whether a comparable biological or physical system achieves the same function at a dramatically lower resource cost (Rao's squirrel-brain comparison). A multi-order-of-magnitude gap between an engineered system and its biological analog is a strong signal that the engineering approach, not the underlying physics, is the limiting factor.

Prove a novel computational principle in simulation before committing to physical fabrication

Following Rao's sequencing (an open-source simulated model demonstrating the oscillator-based approach, released before the physical chip was built), when pursuing a genuinely novel technical approach with a long and expensive physical validation cycle, look for a way to validate the underlying principle in a cheaper, faster simulated or software form first, to de-risk the physical investment that follows.

Bottom Line

Naveen Rao's argument is that AI's real near-term ceiling is energy rather than chip supply, and that the reason current hardware is roughly ten billion times less energy-efficient than biological brains is a specific, identifiable architectural choice (separating memory and compute, moving vastly more data than necessary) rather than an unavoidable physical limit, which is why his company built a working alternative chip, with memory and compute unified into the same physical element, and validated its efficiency claim with a fabricated prototype within months rather than years.

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free