Flash sale 33% off with code LAUNCH33 Ends in --:--:--
See pricing
All Things PM
Robot-Use Agents: Why General-Purpose Models May Win in Robotics
Y Combinator Startup PodcastAI

Robot-Use Agents: Why General-Purpose Models May Win in Robotics

The founders of Waddle Labs and RoboCurve explain why coding agents are turning into robot controllers, why latency is now the real blocker to real-time robots, and why they think general-purpose robots are about two years away.

September 26, 2026 · 30 min listen · 11 min read · Hanmei, Vincent, Jay
0:00
–:––

Context

MIT professor Philip Isola argued in a recent essay that coding agents are generalizing into robot use agents: general-purpose models that can control different robots with little or no robot-specific training. This episode of Decoded brings on the founders whose demos helped inspire that essay, Hanmei and Vincent of Waddle Labs (building the harness and data pipeline that lets LLMs control robots) and Jay of RoboCurve (an evals company that measures robots, models, and vision-language-action systems across many types of hardware). Together they trace the research path from early robot control papers to today's general-purpose models, and demo a model controlling a robot arm live.

The Big Idea

Coding agents did not just get good at writing software, they got good at writing the code and tool calls that control physical robots, and the same general-purpose model that reasons and writes code well is turning out to also be the best robot controller, not a robot-specific architecture.

The founders believe the field is roughly two years from general-purpose robots that can follow a natural-language instruction and do what a competent teenager could do with their hands, and the biggest remaining blocker is not intelligence, it is latency.

Key Insights

Robot control is becoming a coding-agent problem

  • Vision-language-action models (VLAs) like the 2023 RT2 paper output a robot's next physical pose directly, the same way an early language model just output a final answer with no chain of thought.
  • Newer approaches let the model reason step by step and write code or tool calls that control a robot, the robotics equivalent of chain-of-thought reasoning.
  • The founders argue the real difference between a strong VLA and a strong general-purpose language model is not architecture, it is training approach: a language model pre-trained on web text, images, and code, then lightly adapted, can out-transfer a model trained narrowly on robot data.

Computer-use and CAD data transfer to robots

  • Feeding a model computer-use data, like dragging a cursor to orbit a 3D object in CAD software, teaches it concepts such as top, bottom, left, and right that turn out to be directly useful for controlling a robot arm in physical space.
  • Waddle Labs' founders believe this is a large part of why newer general-purpose models jumped so much on spatial-intelligence and robot-control tasks compared to models from just six months earlier.
  • One founder noted the irony: graphical user interfaces were originally built in the 1980s to make computers feel more like the physical world, and that same data is now what is teaching models to act in the physical world.

In-context learning has a hard ceiling

  • A model can improve on a held-out task purely by being shown more examples in its context window, with no weight updates and no training compute cost.
  • That improvement is not smooth: it can get worse before it gets better, and it saturates after roughly 20 to 40 examples, after which adding more examples stops helping.
  • The absolute ceiling is the model's trained context length: if a model was trained on 100,000 tokens of context, it may only reliably improve up to about half that, and performance degrades again once the window is exceeded because the model can no longer attend over everything in it.

Latency, not intelligence, is now the bottleneck

  • Running a frontier model live inside the control loop for every single robot action is too slow to be economically useful, and the founders demoed this lag directly on a block-picking task.
  • Model latency for this class of model has reportedly been improving at roughly 2x per month; if that trend holds, real-time control could arrive by the end of the year.
  • The near-term fix is not waiting for faster models, it is compiling a model's first attempt at a task into a faster, reusable skill or policy that can run repeatedly without a large model reasoning at every step.

Code as policies turns skills into reusable tools

  • The 2022 "code as policies" research from Google DeepMind gave a coding agent a fixed list of robot primitives, like pick up an object or move to a pose, as literal Python functions, and let it write code combining them to do complex tasks in one shot, with no extra robot-specific training data.
  • This built on the earlier Voyager approach in Minecraft, where a coding agent used built-in tools to assemble new tools on the fly and reused them later, compressing experience into a callable skill.
  • In Waddle Labs' current harness, repetitive, well-understood parts of a task get compiled into deterministic code, while a vision-language model is only inserted at the points that genuinely vary, such as detecting an object or deciding what to do after a failure.

General-purpose robots are roughly two years out

  • There is reported consensus among frontier AI labs and robotics foundation-model companies that general-purpose robots, ones that can follow any natural-language instruction and do what a competent teenager could do by hand, are roughly two years away or sooner.
  • The founders compare this to the ChatGPT moment for robotics: generalizing to tasks and environments the model was never specifically trained on, rather than needing a bespoke dataset per task.
  • They believe society and much of the market are not yet prepared for how fast this capability is arriving.

Mental Models & Frameworks

Transduction versus program induction

  • Transduction means training a function that maps inputs directly to outputs, which is information-inefficient: getting from X to Y with few examples requires a lot of built-in assumptions about the problem.
  • Program induction instead asks a model to emit a function (a program) that performs the mapping, using a handful of example input-output pairs the way a coding interview gives a few examples before asking for working code.
  • The founders frame today's coding agents as bridging this gap: earlier symbolic program-synthesis research struggled because finding the right built-in assumptions was hard by hand, while a strong coding model can now write that mapping function directly, described as a "neuro-symbolic" approach where neurons emit symbols.

The learning consolidation hierarchy

  • In-context learning (appending new examples straight into the prompt) is the cheapest and fastest way to improve a model on a task, but it caps out quickly and vanishes once the context window is exceeded.
  • Beyond that ceiling, the next options in order of cost are retrieval into active memory (pulling in similar past examples on demand), then LoRA fine-tuning at increasing rank, then full supervised fine-tuning or reinforcement learning on the base weights.
  • The founders frame a live open question as how to compress what a model learns in-context (through repeated attempts) down into a faster tool, skill, or eventually an updated set of weights, rather than re-deriving it in context every time, the way sleep is thought to compress a day's experience into long-term memory in humans.

The Platonic Representation Hypothesis

  • Borrowed from Philip Isola's paper of the same name (a reference to Plato's Cave), the idea is that models trained on very large, different datasets under different objectives tend to converge on similar internal representations of the world.
  • Applied to robotics, this predicts that a very strong general-purpose language model will end up with representations similar to a very strong robotics-specific model, so one sufficiently strong general model can outperform several narrower, specialized ones.
  • The founders treat this as the strongest version of the "bitter lesson," that general methods with more data and compute beat hand-designed, domain-specific ones, and use it to explain why a single frontier model can already out-transfer purpose-built robot models.

Trade-offs & Nuance

Direct model control versus a compiled harness

  • Calling a frontier model directly for every single robot action (like Astra controlling a robot arm live) is flexible and needs no extra engineering, but the founders showed it is currently slow, bottlenecked by the model's own response latency.
  • Routing repetitive or previously-solved tasks through a compiled harness or skill library instead runs faster and handles edge cases more reliably, but only works well once that skill has actually been learned from a prior attempt.
  • The practical answer isn't choosing one approach permanently: use the large model directly to explore a new task, then compile what worked into faster reusable code once the task is understood, keeping the model in the loop only at points of genuine variation.

Practical Application

Reserve the large model for real variation only

  • When building an agent that controls a physical or digital system repeatedly, do not put the frontier model in the loop for every single step.
  • Identify which parts of the task are deterministic and repetitive (approaching an object, a fixed sequence of moves) and compile those into ordinary code.
  • Call the model only at the points where a decision genuinely varies, such as classifying what object is in view or deciding how to recover from a failure.

Compile successful attempts into reusable skills

  • The first time an agent solves a new task, treat that trace as raw material, not just a completed job.
  • Convert the working sequence into a callable skill or policy that can run again without the large model reasoning through every step from scratch.
  • This is what let Waddle Labs move from a slow, model-in-the-loop demo to something that could plausibly run at production speed later.

Test data transfer, not just in-domain accuracy

  • Before assuming a model needs domain-specific training data, test whether data from an adjacent domain, like computer-use or CAD interaction data, already improves its performance on the target task.
  • The founders found spatial-reasoning gains from data that was never intended for robotics, which suggests the cheapest performance win may be broader pre-training data, not more narrow, expensive robot-specific data collection.

Know where in-context learning will stop helping

  • If you're building a product that leans on a model improving from examples given in the prompt, expect gains to plateau after roughly 20 to 40 examples and to degrade once you approach the model's trained context length.
  • Design the product to hand off to retrieval, fine-tuning, or a compiled skill once that ceiling is reached, rather than assuming more examples in context will keep helping indefinitely.

Questions to Consider

  • Where in our own product does a model reason through every step from scratch, when a past successful run could instead be compiled into a faster, reusable tool or skill?
  • Are we assuming a capability requires narrow, domain-specific training data, when data from an adjacent domain might already transfer, the way computer-use and CAD data reportedly improved robot spatial reasoning here?
  • If our product relies on in-context learning to improve a model's output as a user adds examples, do we know where that improvement plateaus, and what happens to the product once a user exceeds it?
  • Is latency, not model intelligence, actually the constraint holding back a feature we've shelved as "not good enough yet"?

Bottom Line

The same general-purpose coding and reasoning models that write software are turning out to be strong robot controllers too, and the founders behind this shift believe general-purpose robots that follow plain-language instructions are roughly two years away, with latency, not intelligence, as the main remaining barrier.

Whatever system you're building on a frontier model, check whether the real blocker is the model's capability or simply how often you're forced to call it live instead of compiling what it already learned.

Concepts to Explore

The bitter lesson

  • The idea, from a well-known essay by Rich Sutton, that general methods leveraging more data and compute eventually beat hand-crafted, domain-specific approaches.
  • The founders invoke it directly: rather than building a robot-specific architecture, they'd rather build one very strong general-purpose model and use it to control robots, betting that scale and general data beat specialization.

Code as policies

  • A research approach where a coding agent is given a robot's basic actions as literal callable functions and writes code combining them to complete a task, without needing task-specific training data.
  • It matters here because it is the direct ancestor of today's harnesses, which mix deterministic compiled code with a model called in only at points of real uncertainty.

Resources Mentioned

ResourceTypeWhy it was mentioned
RT2 paperResearch paperCited as one of the earliest successful uses of a pre-trained language model to directly output robot control poses.
Code as Policies (Google DeepMind, 2022)Research paperDiscussed in depth as the origin of using a coding agent's callable functions to control robots in one shot.
VoyagerResearch paperCited as an earlier example, in Minecraft, of a coding agent creating and reusing its own tools from built-in primitives.
The Platonic Representation Hypothesis (Philip Isola)Research paperThe framework used to explain why one strong general-purpose model can match or beat specialized robotics models.
On the Measure of Intelligence (Francois Chollet)Research paperCited as the source of the transduction versus program-induction distinction used to describe how coding agents solve robot control.

People to Follow

Philip Isola

MIT professor whose essay on robot use agents, and earlier paper on the Platonic Representation Hypothesis, frame much of this episode's argument that general-purpose models are converging toward similar, transferable representations of the world.

Francois Chollet

Researcher known for the ARC benchmark and the transduction-versus-program-induction framing the founders use to explain why coding agents can act as robot policies with very little robot-specific data.

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free