Context
MTS host Sophia Dew speaks with Lukasz Kaiser, co-author of the landmark "Attention Is All You Need" paper that introduced the transformer architecture underlying modern AI, at the Open Source AI Summit in San Francisco. The conversation asks a direct question: is AI power concentrating in a handful of large companies because that is an inherent feature of the technology, or because it is simply the cheapest path with today's architecture. Kaiser, who helped invent the transformer and later worked at OpenAI, argues concentration is a property of this specific technological moment, not a permanent state, and explains what would need to change for smaller players and open research to compete again.
The Big Idea
AI power is concentrating today because the transformer architecture rewards scale (more data, more compute, more money) rather than because concentration is fundamental to how intelligence works, and a research breakthrough toward smaller, more specialized, more efficient models could redistribute that power just as quickly.
Kaiser points to the fact that transformers are less than a decade old and argues that humans, who are not generalists trained on the entire internet, are living proof that a fundamentally different and more efficient path to intelligence exists, even if nobody has fully found it yet.
Key Insights
Concentration comes from the architecture, not AI itself
Kaiser explains that transformers need enormous, expensive data centers and vast amounts of scraped internet data to get smarter, which naturally favors companies with billions of dollars to spend. He is careful to frame this as a property of the current technology rather than something inevitable: without a research breakthrough that makes models smarter with less data, the only lever companies have is to keep going bigger, and going bigger requires exactly the kind of capital concentration we see today.
Transformers generalize badly outside broad training
Kaiser notes a specific technical limitation: transformers work well when trained on something close to the entire internet, producing a reasonably capable general model, but perform poorly when you try to train them narrowly on just one specific thing. He says this narrow-training weakness is itself an unsolved research problem, not an inherent limit of machine intelligence, and that solving it is part of what would let smaller, specialized models compete with giant generalist ones.
Research labs have shifted from research to product
Kaiser observed this shift directly, having joined OpenAI when it was primarily a research lab. He says it has since become much more of a product company, which necessarily reduces its focus on fundamental research. He frames this as an opportunity rather than only a loss: the space that big labs are vacating on pure research is now open to academia, universities, and independent researchers to fill.
Consumer GPUs now rival the hardware that built the transformer
Kaiser makes a concrete point about how much compute has become available to individuals: he recently bought a single RTX 5090 GPU that has more raw power than the eight-GPU machine his team used to do the original transformer research. He is clear that a single consumer GPU still cannot train a large frontier model, but says it is enough to meaningfully research and experiment, which he argues more people should be doing given how much has become accessible outside big labs.
Mental Models & Frameworks
Ensembles of specialized models over one giant generalist
Kaiser's model for where AI may be headed: instead of one enormous model trained on all of the internet, a large number of smaller, specialized models, each strong within its own domain, could collectively learn more efficiently from a fixed amount of data than a single generalist model does today. He points to humans as existing proof of this pattern: no single human brain contains all of humanity's expertise, yet distributed human expertise collectively outperforms any single generalist. He references basic research on model ensembles as an early technical hint that this direction is viable, while acknowledging nobody has fully cracked how to make it work in practice yet.
Practical Application
Do not assume today's concentration is permanent
If your product strategy currently assumes only a handful of large labs will ever be capable enough to build on, reconsider on a longer time horizon. Kaiser's argument is that the current advantage of scale is tied to a specific, less-than-decade-old architecture, not a law of AI, so a research shift toward smaller specialized models could change who can credibly compete within a normal planning horizon.
Treat cheap consumer compute as a real research option
If your team has been assuming that meaningful AI experimentation requires large-scale data center budgets, Kaiser's example of an RTX 5090 outperforming his original transformer-research hardware suggests otherwise for research and experimentation, even though it is not enough to train a frontier-scale model outright. Smaller teams and individual researchers now have access to more raw compute than existed when the technology that runs today's largest models was first invented.
Questions to Consider
- If a research breakthrough made smaller, specialized models competitive with today's largest general-purpose transformer models, how would that change your team's build-versus-buy decisions for AI features?
- Is your organization treating today's AI landscape, where a handful of large labs dominate frontier capability, as a permanent constraint, or as one stage of a technology that Lukasz Kaiser says is less than ten years old?
- Where in your product could a smaller model, trained narrowly on your own domain data rather than the whole internet, potentially outperform a giant general-purpose model if the narrow-training weakness Kaiser describes gets solved?
Bottom Line
Lukasz Kaiser argues today's AI power concentration is a byproduct of the transformer architecture's hunger for scale, not a fixed feature of intelligence itself. He sees real signs, cheap consumer compute and labs shifting away from pure research, that open the door for smaller players and independent researchers to help find the next breakthrough that could redistribute AI capability more broadly.
People to Follow
Lukasz Kaiser
Co-author of "Attention Is All You Need," the 2017 paper that introduced the transformer architecture underlying nearly all modern large language models. Kaiser later worked at OpenAI during its transition from a primarily research-focused lab to a major product company, and now researches independently, including hands-on experimentation with consumer-grade GPU hardware.
