All Things PM
What the Top AI Users Are Doing Differently
The AI Daily Brief: Artificial Intelligence News and AnalysisAI

What the Top AI Users Are Doing Differently

New OpenAI research shows the gap between average and frontier AI users has grown from 2.6x to 8.3x in months, and it's not about better models, it's about agents, skills, and a handful of specific usage patterns any team can copy.

August 25, 2026 · 28 min listen · 7 min read
0:00
–:––

Context

NLW opens with a headlines roundup (Meta's upcoming consumer agent, price cuts from OpenAI and GrokBot, Hugging Face's acquisition rumors, Nvidia's investment spree, and a Microsoft internal spreadsheet on employee AI spend) before turning to the main segment: new OpenAI research on how frontier enterprises use AI differently than everyone else. The core finding is that the usage gap between the most advanced enterprise AI users and average ones has exploded in just months, driven almost entirely by the shift from chat-based AI to agentic AI. For a PM, this episode is less about any single feature and more a diagnostic checklist: which specific behaviors separate a team compounding its AI advantage from one that's flat.

The Big Idea

The gap between frontier AI users and everyone else isn't about which model they use, it's about how much of their work has shifted from chat-style assistance to agentic delegation, and that shift compounds: the farther ahead a firm gets, the faster it pulls away.

OpenAI's own data shows the gap between frontier firms (the top 10% of enterprise users by usage) and typical firms grew from 2.6x in January to 8.3x by June, with frontier firms now producing 17 times more output tokens than they did a year and a half ago, versus roughly 2 times for average firms.

Key Insights

The frontier-to-average usage gap tripled in five months

OpenAI defines a "frontier firm" as one in the top 10% of usage in a given month, measured by output tokens per active user, versus a "typical firm" in the 45th-55th percentile. From roughly April 2025 through most of the year, frontier-firm users produced about 2x the output tokens of typical-firm users, a gap that held steady. Starting around October it began widening, reaching 2.6x by January, then exploding to 8.3x by June. Over that same longer period, frontier firms increased their total output tokens 17x while typical firms only doubled theirs. NLW frames this as evidence that agentic use doesn't just add value, it compounds a lead once a firm gets ahead.

Agentic work overtook chat work inside a year

OpenAI tracked the balance between ChatGPT-style chat tokens and agentic tokens (a proxy for delegated, multi-step work) across enterprise usage:

  • August 2025 (GPT-5 launch): essentially 100% chat tokens, 0% agentic.
  • February 2026 (Codex app for Mac launches): 87% chat, 13% agentic.
  • March-April 2026 (GPT-5.4, Codex for Windows): 73% chat, 27% agentic.
  • Late April 2026 (GPT-5.5): the "flippening," agentic crosses 53%.
  • June 2026 (data cutoff): 36% chat, 64% agentic.

The practical read: if output tokens are a proxy for how much actual work gets done, roughly two-thirds of enterprise AI-driven work is now agentic rather than conversational, a complete reversal in under a year.

Non-technical roles are adopting agents fastest

A widely shared OpenAI chart tracked Codex user growth by job title, indexed from February 2026. Engineering and technical roles grew 5x, the smallest increase of any function, while non-technical roles grew far faster from a lower base: finance and accounting up 20x, marketing and communications up 26x, people and recruiting up 41x, sales and account management up 41x, and legal up 108x. NLW notes this cuts against the assumption that coding agents stay confined to engineers, agentic coding tools are increasingly being picked up by non-engineers to build things inside their own function.

Skill and plugin adoption predicts who pulls ahead

OpenAI measured what share of weekly active users engage with plugins (connectors to other apps or data) or skills (reusable, saved workflow instructions). At typical firms, about 9% of users use plugins and 3% use skills. At frontier firms, that's 21% using plugins and 19% using skills, roughly 2.3x and 6x higher respectively. Internally at OpenAI itself, 95% of employees use plugins and 93% use skills, suggesting even frontier enterprise customers still have significant headroom before hitting the ceiling OpenAI's own team has reached.

Comparing how legal teams use chat versus agentic tools illustrates the broader shift in work type, not just volume. In chat mode, writing makes up 57% of legal usage and knowledge retrieval 20.5%, with system operations at a negligible 0.2%. In agentic mode, writing drops to 16.2% and knowledge retrieval to 8.3%, while system operations jumps to 17.7%, workflow automation to 7.7%, classification and extraction to 4.6%, and coding (building small applications, despite these being non-engineers) rises to 32.9%. NLW notes this same pattern, chat staying concentrated in writing and retrieval while agentic work shifts into systems integration, shows up across every department OpenAI examined.

Mental Models & Frameworks

The agentic use-case ladder

NLW describes a four-rung progression for how agentic work matures within a function, useful for diagnosing how advanced a team's AI usage actually is versus how advanced it feels:

  • Generation: producing a single artifact, like drafting an email, a report, or an Excel formula.
  • Synthesis: pulling from multiple disparate data sources to produce something more complete than any one source alone.
  • Execution: the AI actually interacts with existing systems and takes action within them, not just producing text about them.
  • Maintenance: the agent doesn't just execute a task once but maintains a system continuously over time.

In the legal example from the episode, this maps to agents comparing contract terms, flagging deviations, drafting redlines, and monitoring ongoing commitments (execution and maintenance), while humans retain negotiation of material terms, risk tolerance, and final approval (judgment and accountability). Use this ladder to evaluate whether a team's "agentic AI usage" is actually generation dressed up in agent tooling, or genuine execution and maintenance work.

Single-player versus multiplayer agentic work

NLW argues that even at frontier firms, most agentic use still happens inside individual silos, one person's agent doing one person's task. He predicts the next major gains will come from "multiplayer" or team-level AI that operates at the intersection of multiple people's or teams' work, with shared context and handoffs, rather than each person running their own isolated agent. This reframes the adoption question from "how many of my people use agents" to "do our agents actually coordinate with each other."

Practical Application

Track output-token share, not just tool logins

Measuring whether a team is using AI meaningfully requires looking at how much of the actual work output is agent-driven versus chat-driven, not just adoption or login rates. A team stuck at high chat usage but near-zero agentic usage looks identical to a "successful" AI rollout on a login dashboard while sitting nowhere near the compounding gains frontier firms are capturing.

Push adoption into skills and plugins specifically

Given the 2-6x adoption gap in skills and plugins between typical and frontier firms, prioritize getting a critical mass of the team actually building and reusing skills and connecting plugins, not just prompting a general assistant. This is a more specific, measurable target than a vague "increase AI usage" goal.

Audit non-engineering functions for agent-shaped work first

Since non-technical roles (legal, sales, marketing, recruiting) are growing agentic usage fastest despite starting from the lowest base, look for agent-shaped work outside engineering rather than assuming coding agents only matter to developers. A legal, sales ops, or recruiting workflow with repeatable steps and available context is a stronger early candidate than it might initially seem.

Map a team's actual position on the use-case ladder

For any function using AI today, explicitly identify whether its usage sits at generation, synthesis, execution, or maintenance. A team assuming it's "doing agents" because it drafts emails with AI is still at the bottom rung; moving up the ladder, not just increasing volume at the same rung, is what produces the compounding gap frontier firms are seeing.

Bottom Line

The 2.6x-to-8.3x usage gap between frontier and average AI users isn't about access to better models, every firm has access to the same models, it's about the shift from chat-based assistance to agentic delegation, and that gap compounds the longer a firm waits to make the shift.