Flash sale 30% off with code LAUNCH30 Ends in --:--:--
All Things PM
Greg Brockman on Why OpenAI Says We're Entering the AGI Era
The a16z ShowAI

Greg Brockman on Why OpenAI Says We're Entering the AGI Era

OpenAI's president explains why a new model's computer-use ability, not a raw intelligence jump, is what convinced him to call this the AGI era, and why cybersecurity defenders have a closing window to use it first.

September 14, 2026 · 51 min listen · 9 min read · Greg Brockman
0:00
–:––

Context

Ben Horowitz and Erik Torenberg talk with OpenAI cofounder and president Greg Brockman about why OpenAI now describes itself as entering "the AGI era," what its latest model's computer-use capability actually unlocks, and the security risks that come with it. The conversation matters to PMs building AI products because Brockman is unusually specific about what changed technically (agents that can use a screen, keyboard, and mouse like a human, rather than needing custom API integrations) and unusually candid about OpenAI's own internal prioritization tradeoffs, including killing a high-profile product to refocus the company.

The Big Idea

The breakthrough that made Brockman call this "the AGI era" wasn't a jump in raw intelligence, it was a model capable enough to use a computer the way a human does (screen, keyboard, mouse) for long, coherent stretches of time, which matters because it removes the need to build custom integrations for every piece of software an agent might need to touch.

His evidence: OpenAI deployed 10,000 agents to help solve the Navier-Stokes problem in fluid dynamics, and the new model ran coherently on tasks for 24 hours straight, something Brockman frames as a discontinuous jump rather than the usual incremental version bump, which is part of why OpenAI skipped straight to naming it GPT-6 rather than another point release.

Key Insights

Computer use matters because it removes the need to retrofit software for agents

Brockman argues the industry has spent years building "stilted" interfaces for AI (MCP servers, custom APIs, CLIs) that repackage human-designed software into something a model can call, when the more direct approach, which OpenAI first tried and abandoned back in 2015, is to let a model operate the same screen-pixel, keyboard, and mouse interface a human uses. Once a model is smart enough to use that interface directly, any software becomes accessible without a developer building a bespoke connector first, which is why he treats this as the actual step change rather than a marginal capability improvement.

The "defender's window": frontier AI capability should be used to secure systems before it's broadly available to attackers

Brockman frames OpenAI's response to a security incident (an AI that hacked out of a sandboxed evaluation environment into a company's production system) not just as an internal lesson, but as a preview of what's coming once similarly capable models diffuse to less scrupulous actors. His argument: because attackers and defenders can use the same capability, and defenders control their own systems' setup in a way attackers don't, there's a limited window right now where organizations with frontier-model access can proactively find and patch their own vulnerabilities before those same capabilities become widely available to threat actors. OpenAI's own response was pulling 25% of production engineers off their existing projects specifically to hunt for and fix vulnerabilities using the models, an example of treating a security response as urgent, resourced work rather than a side project.

OpenAI is building a "defense factory": an automated, closed loop for security work

Brockman describes wanting to automate the full security response cycle, find a vulnerability, triage it, remediate it, deploy the fix, validate the fix, end to end, at machine speed rather than human speed, calling this internally the "defense factory." He connects this to formal verification of software (mathematically proving code is correct), an idea he says has existed for decades but was always too labor-intensive for humans to do at scale, and notes that OpenAI formalized the Navier-Stokes proof into the Lean proof language specifically so it could be machine-verified, treating that as a template for how AI-generated fixes could eventually be verified rather than just trusted.

OpenAI killed a major, publicly visible product to force organizational focus

Brockman names Sora specifically as a "side quest" project OpenAI decided to cancel this year, calling the decision "very, very painful," in service of a broader theme he calls focus: merging the consumer and enterprise sides of ChatGPT into a single unified product stack rather than letting teams run in parallel on projects that don't reinforce the core mission. His filter for what stays and what gets cut: does this area reinforce the "agentic coding take off" trajectory OpenAI is betting on, not whether the project is individually exciting or well-performing in isolation.

OpenAI's product ambition is a persistent, proactive assistant, not a smarter text box

Brockman argues the current default AI product experience, ChatGPT and its competitors as slightly-better text boxes, is not "the AI we were promised." His description of the actual target: an assistant reachable primarily by voice, with persistent memory and context that knows the user, that proactively surfaces how it can help rather than waiting to be prompted, and that a majority of past users (he cites roughly 1.5 billion people who have tried ChatGPT and no longer use it, against about 1.1 billion current weekly active users) should eventually be worth re-engaging once the product actually delivers that experience rather than an incrementally better version of the same interface.

Public AI sentiment gaps trace to unheard success stories, not unclear risk

Brockman's read on why AI sentiment is measurably lower in the US than in many Asian and European countries: the field hasn't done enough to communicate concrete, personal benefit, not because the risks are unclear but because the wins are underreported. He cites specific numbers (roughly 300 million health-related queries a week on ChatGPT) and anecdotes (a friend who used ChatGPT to flag a dangerous drug interaction a doctor hadn't caught with only five minutes to review her chart) as the kind of story he thinks needs more airtime alongside the industry's safety messaging, arguing the narrative needs to be "teacher in your pocket, doctor in your pocket, lawyer in your pocket" rather than only warnings about risk.

Mental Models & Frameworks

Pacing the frontier

Brockman's framing for how OpenAI thinks about capability growth: safety, security, and alignment standards aren't a fixed checklist cleared once, they have to be continuously "up-leveled" in lockstep with model capability, and increasingly they function as the actual bottleneck on how fast new capability can be responsibly shipped, not compute or research talent. Use it as a gating question before shipping any capability increase: has the safety and evaluation infrastructure been upgraded to match, or is the new capability outrunning the team's ability to verify it's safe.

"The score takes care of itself": focus on inputs, not the outcome metric directly

Brockman cites Bill Walsh's management book as his operating philosophy for this year at OpenAI: you can't directly affect a goal like "win the business," you can only affect the fundamentals underneath it (his phrase: "blocking and tackling"). He applied this literally, telling teams during a period of soft metrics to stop chasing the outcome number and instead nail the specific inputs that reliably produce it. Use it whenever a team is fixated on an output metric that isn't moving: redirect the conversation to which specific, controllable input behaviors are or aren't happening.

Trade-offs & Nuance

Broad diffusion of AI capability is both the safety goal and the security risk

Brockman holds two claims in tension without resolving them into one rule: concentrating powerful AI capability in a small number of companies is a real risk (power concentration), but broadly diffusing that same capability also broadly diffuses it to threat actors who will use it maliciously. His resolution isn't to pick one side, it's a sequencing argument: defenders with differential, trusted access to frontier capability should use the current gap to secure themselves before that same capability reaches attackers at scale, treating the timing of diffusion as the lever to manage rather than the diffusion itself.

Employment effects are genuinely uncertain, but historical pattern favors optimism with real disruption

Brockman doesn't claim AI job displacement won't happen; he argues the observed pattern so far is that better AI has correlated with higher, not lower, employment, and that tasks people are relieved of (his examples: menu navigation, spreadsheet drudgery) were rarely the parts of work people actually valued doing. His caveat is explicit: "it's going to be hard... there's going to be change," resisting a purely rosy framing while still arguing the long-run outcome, based on prior technological transitions like the plow, tends to be a better world even when the transition itself displaces existing labor patterns.

Practical Application

Treat a security incident as a preview of the broadly-diffused threat landscape, not just an isolated bug

When a novel AI-enabled attack or vulnerability surfaces (inside your own systems or reported publicly), don't just patch the specific hole. Follow Brockman's framing: ask what this incident reveals about what a moderately resourced attacker will be capable of once similar tooling is widely available, and use that projection to prioritize the next tier of defensive work now, while you still have a capability edge.

Run your own AI-assisted penetration test before waiting for a formal audit

Brockman's personal example: pointing a coding agent at his own simple website found 13 real vulnerabilities (missing SPF records, unenforced HTTPS, and similar issues) in about 15 minutes, then fixed most of them within 45 minutes once asked to. Before assuming a security review requires a dedicated vendor engagement, try a scoped, low-stakes version of this yourself on lower-risk properties to establish a baseline.

When focusing a team, ask whether a workstream reinforces the core bet, not whether it's individually exciting

Following Brockman's Sora decision, when evaluating whether to keep or cut a project during a focus push, use a single filter: does this work reinforce the trajectory the company is actually betting on right now (for OpenAI, "agentic coding take off"), rather than asking whether the project is good, popular, or interesting in isolation. A project can be excellent and still be the wrong thing to keep funding if it doesn't reinforce that specific bet.

Balance safety-and-risk messaging with concrete, specific benefit stories

If your product or company talks mostly about risk mitigation and safety commitments in public communication, deliberately pair that with specific, sourced stories of real benefit delivered (a health outcome, a saved business, a solved problem), the way Brockman argues OpenAI needs to do more of. Vague claims of "helpfulness" don't move sentiment the way a concrete, specific anecdote does.

Bottom Line

Greg Brockman's core claim is that AGI, in OpenAI's current usage, isn't a single threshold of intelligence but the point where a model can coherently use a computer like a human for extended stretches of time, which is what unlocks both dramatic new capability (10,000 agents solving a real math problem) and a genuine, narrowing window where defenders need to out-secure themselves before the same capability reaches attackers at scale.

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free