All Things PM
The AI Model Tier List
The AI Daily Brief: Artificial Intelligence News and AnalysisAI

The AI Model Tier List

A viral AI model tier list and a widely shared enterprise-spend chart both got read as proof one model had "lost," but the real story is that businesses are quietly assembling multi-model stacks, and the chart everyone shared was measuring something else entirely.

August 24, 2026 · 29 min listen · 13 min read
0:00
–:––

Context

NLW breaks down two related stories: a viral AI model tier list from content creator Theo ranking today's leading models, and a widely shared Financial Times chart implying that Anthropic's flagship model had failed to gain enterprise adoption because of its price. Both are used to make a broader point: the era of a single "best" model mattering above everything else is ending, replaced by businesses assembling model stacks that mix premium and open-weight models for different tasks, costs, and speeds. The episode also covers Hugging Face reportedly exploring a $13 billion sale, Nvidia's deepening bets on open models through deals with Poolside, Mercor, and Perplexity, and AT&T's active shift toward open models to cut AI costs. For a PM building or buying AI-powered products, the throughline is that model selection is becoming a genuine architecture decision, not a one-time pick, and that popular narratives about which model is "winning" often rest on data that doesn't mean what it appears to.

The Big Idea

The AI industry has moved past the question of which single model is best; the real competitive and product decision now is how to architect a stack of multiple models, open and closed, matched to specific tasks by cost, speed, and capability, and popular narratives claiming one model "lost" often reflect data artifacts rather than genuine capability verdicts.

NLW's clearest evidence for this: a viral chart implying enterprises had rejected Anthropic's Fable 5 over price turned out to be explained mostly by a mandatory 30-day data retention policy tied to a government safety review, not by businesses judging the model unworthy of its cost.

Key Insights

A viral chart's real story got missed

  • What happened: a Ramp AI Index chart, shared by Ramp's lead economist Eric Arazi, showed enterprise spend heavily favoring Anthropic's Opus models over Fable 5, and was picked up by the Financial Times with the framing that businesses "don't think [Fable 5 is] worth the price."
  • What was missed: Fable 5 shipped with a mandatory 30-day data retention policy, a provisional safeguard tied to a government safety review after the model was briefly banned, which alone is enough to disqualify it for many enterprises regardless of price or capability.
  • A second confound: the data comes from Ramp's token-and-spend-management product, whose users are self-selected for cost-consciousness, drawn from an extremely tech-forward customer base that isn't representative of the broader enterprise market, which moves far more slowly.
  • Why it matters for a PM: a chart with a clean, dramatic headline can miss a compliance or policy variable that fully explains the pattern; before trusting a data-driven narrative, including your own product's usage data, check what population generated it and what non-obvious variables could explain the pattern before concluding it's a verdict on quality.

Businesses are assembling multi-model stacks

  • The pattern: AT&T is deliberately holding its OpenAI and Anthropic spend flat while growing its use of open models, Nvidia's Nemotron plus models from Meta and Google, to handle a growing share of internal AI queries, currently 40%, targeting 60 to 70% over the coming years.
  • Where the line still holds: AT&T still routes advanced tasks like generating code to frontier closed models, and only shifts simpler tasks, like summarizing a pull request, to open models.
  • The router payoff: for AI coding specifically, AT&T's Vice President of Data Science Mark Austin reported that using a model router cut costs by as much as 56% while quality dropped only about 2%.
  • Why it matters: this is a live example of the model-stack strategy the episode centers on, treating model choice as a per-task routing decision rather than a single company-wide default, and it shows the efficiency gain is real and measurable, not theoretical.

Enterprise AI adoption moves very slowly

NLW pushes back on the assumption embedded in a lot of tier-list and model-comparison discourse, that it's shocking when enterprises haven't adopted a model that's only a few months old, by pointing out that he regularly hears from people still using models from nine months ago simply because that's what their company has provisioned. The practical implication is that any usage or adoption data reflecting real enterprise behavior is going to lag genuine capability rankings by a wide, structural margin, and reading that lag as a verdict on model quality is a category error.

The highest-rated model isn't the default

Theo, the creator of the viral tier list, rated Anthropic's Fable 5 the only model in S tier, above GPT-5.6 Sol's A tier, yet said he would still default to Sol day to day and would miss it more if it were gone. His reasoning: Fable 5 is unbelievably capable but "trips over things," takes unnecessary shortcuts, and occasionally loses track of what it's doing, so he reserves it for code he actually wants to merge, double-checking other models' work, and hard conceptual conversations, while treating Sol as the faster, more predictable model he initiates most tasks with. The lesson generalizes beyond this one comparison: a model's ceiling on raw capability and its fit for a team's actual day-to-day workflow are two different rankings, and optimizing around the first while ignoring the second produces a worse outcome than picking the model that's merely very good but reliable.

Nvidia is buying into model training

Nvidia made three separate moves in one stretch that point toward the same strategy: participating in data-labeling startup Mercor's funding round (Nvidia already uses Mercor for reinforcement learning on its Nemotron models), taking a stake in Perplexity's new $30 billion round, and striking a $6 billion non-exclusive licensing deal plus a $1 billion equity investment in Poolside, a coding-focused open model startup, which also brought over 100 Poolside engineers, described as the bulk of its engineering team, into Nvidia to build what the plan calls the world's most powerful open model, aimed at rivaling Chinese labs DeepSeek and Moonshot. Nvidia's own hardware business benefits either way: more open-model activity, wherever it happens, still runs on GPUs, so Nvidia is directly funding the side of the model market most likely to keep chip demand broad rather than concentrated in a handful of closed-model labs it doesn't control.

Nvidia is raising chip prices sharply

Nvidia has begun notifying customers that top-end Grace Blackwell and Vera Rubin chip prices will rise by as much as 17%, applying even to orders already placed for delivery next year; a full 72-chip Vera Rubin rack is expected to reach $8 million, adding roughly $5 billion to the cost of building a gigawatt of compute. Reporting attributes the increase to spiraling memory costs, and one source said cloud providers will almost certainly pass the increase on to the customers renting their chips. For any team budgeting AI infrastructure spend for next year, this is a concrete signal to build in headroom for compute costs rising independent of usage growth.

Mental Models & Frameworks

The three-tier model spend framework

MIT's Christian Catalini splits AI spend into three categories:

  • Cheap generalist: commodity open-weight models, used for high-volume, lower-stakes tasks where "good enough" is fine.
  • State-of-the-art generalist: frontier closed-lab tokens, the premium tier, used when raw capability ceiling matters more than cost.
  • State-of-the-art specialist: models that combine open weights with a company's own proprietary context through post-training, aimed at frontier-level performance on a narrower, enterprise-specific task set for less than the pure frontier price.

Use it to plan an AI product's model architecture deliberately across these three tiers rather than defaulting to one vendor for everything. Microsoft's Foundry product, which lets companies post-train Microsoft's base models on their own data, is built around exactly the third category.

Closed models keep the economic value

Investor Gavin Baker predicts open-source models will carry the majority of token volume, he estimates closed frontier tokens will end up as only 15 to 25% of total tokens, while still capturing 60 to 90% of total economic value, because premium intelligence commands a premium price even while handling a minority of raw task volume. Use this as a model for planning revenue or value around a model stack: a "high volume, low margin" open tier and a "low volume, high margin" frontier tier can coexist, and each contributes differently to total economics, so don't judge a model tier's importance by token share alone.

Trade-offs & Nuance

Open models save cost but add work

Open, cheaper models can meaningfully cut AI costs, AT&T cut coding-task costs by 56% using a model router with only a 2% quality drop, but the savings aren't free or universal. Building and tuning a router, evaluating which tasks tolerate a small quality loss versus which don't, and continuously re-testing as new model versions ship all take real engineering investment, which is why AT&T still routes advanced code generation to frontier closed models rather than shifting everything to open weights. The trade-off is between the ongoing complexity of maintaining a multi-model architecture and the cost savings it enables; for a team without dedicated capacity to build and maintain that routing layer, defaulting to a single frontier model may be the better trade even at a higher token cost.

Practical Application

Audit your own usage-data narratives

Before trusting or presenting a chart of your own product's or team's usage data as a verdict on quality, cost, or preference, check who generated the data, a self-selected, cost-conscious segment of users, a specific product's customers, and what policy, compliance, or configuration variable could fully explain the pattern before concluding it reflects a genuine capability or preference judgment. This is exactly the mistake in the widely shared Ramp and Financial Times chart on Anthropic's Fable 5, where a 30-day data retention policy tied to a government safety review, not price or quality, explained most of the low enterprise usage.

Design a model router for your product

  • Do: classify the tasks your product routes to an LLM by risk and complexity, is the output reversible, does it need to be merged as code, is it a simple summarization, and route the low-risk, high-volume ones to a cheaper or open model.
  • Then: measure the actual quality delta, not the theoretical one, the way AT&T measured a 2% quality drop against a 56% cost reduction for AI coding tasks, before deciding whether the trade-off is worth it for each task category.
  • Why it works: most products don't need frontier intelligence for every call; treating model choice as a per-task routing decision rather than a single default captures most of the cost savings while reserving the most capable, and expensive, model for the tasks that actually need it.

Separate a model's ceiling from its fit

Before standardizing your team or product on "the best" model by benchmark score, run your own actual workflows through the top two or three candidates and track which one you or your team naturally default to day-to-day, the way Theo rated Fable 5 the single highest-tier model yet said he'd still default to and miss Sol more. A model's raw capability ceiling and its practical reliability for your specific, repeated tasks are different rankings, and the second one usually matters more for day-to-day product decisions.

Budget for compute costs rising anyway

When planning next year's AI infrastructure spend, build in headroom for hardware price increases independent of your own usage growth. Nvidia has already begun raising prices on top-end chips by as much as 17%, even on orders already placed, driven by rising memory costs. A budget that only scales with projected usage growth will likely undercount actual spend if compute costs are rising underneath it.

Questions to Consider

  • Do we have a usage or spend chart in our own product or team that we've treated as a clean verdict on quality or preference, without checking whether a policy, compliance, or configuration variable, like Fable 5's data retention requirement, could fully explain the pattern instead?
  • Are we treating model or vendor selection as a one-time decision, when routing different tasks to different models by cost, speed, and risk, the way AT&T built a router that cut AI coding costs by 56% for a 2% quality loss, could capture savings we're currently leaving on the table?
  • If we ranked the AI models or tools we use by raw capability, would the top-ranked one actually be the one we default to day-to-day, or is there a gap between capability ceiling and practical day-to-day fit, the way the tier-list creator rated Fable 5 highest but still defaults to Sol, that we haven't accounted for?
  • Is our infrastructure or vendor cost budget for next year assuming flat unit prices, when Nvidia has already signaled top-end chip prices could rise by as much as 17% independent of how much more compute we actually use?

Bottom Line

The single-best-model era is ending: businesses, AT&T is a concrete example, are building model stacks that route different tasks to different models by cost, speed, and risk, and the popular charts and tier lists used to declare a winner or loser often reflect selection bias, policy constraints, or a mismatch between raw capability and practical day-to-day fit rather than a clean capability verdict. Treat any model comparison, including your own usage data, with the same scrutiny you'd want applied to a product metric before you act on it.

Tools & Products

Tool / ProductWhat it doesWhy it was mentioned
Ramp AI IndexEconomic analysis of AI spend patterns built from Ramp's token-and-spend-management productSource of the widely misread chart on Anthropic Fable 5's enterprise adoption discussed in Key Insights.
Vercel AI GatewayA routing product for developers to move traffic between multiple AI modelsIts usage data, open-model token share nearly doubling from 28% to 62% in two months, was cited as evidence of the shift toward open models, with the caveat that it's biased toward developer usage rather than the general enterprise market.
Microsoft FoundryLets companies post-train Microsoft's base models on their own proprietary dataCited as the product embodiment of the "state-of-the-art specialist" tier in Christian Catalini's three-way model-spend framework.
OpenRouterA model-routing service recently acquired by Stripe for $7 billionCited as evidence of investor appetite for the router layer that lets businesses move between models rather than commit to one.

Notable Quotes

"Fable is that genius at the company that no one wants to work with, but no one wants to fire because they're the smartest person there. If you learn how to work with them, it's incredible." (Theo)

"This is starting to look like arguing about your favorite color or your favorite Pokemon. In a few years, we will laugh about all this." (Feddy's Intern)