Flash sale 33% off with code LAUNCH33 Ends in --:--:--
See pricing
All Things PM
The Most Important New AI Tools from OpenAI DevDay
The AI Daily Brief: Artificial Intelligence News and AnalysisAI

The Most Important New AI Tools from OpenAI DevDay

OpenAI shipped more than 20 launches at Dev Day, from always-on Dots agents to a shared Space workspace and a model near its top tier for a fifth of the price. NLW argues the real story is three trends getting confirmed, not anything blisteringly new.

September 30, 2026 · 23 min listen · 12 min read
0:00
–:––

Context

NLW (Nathaniel Whittemore) hosts The AI Daily Brief, a daily show on AI news. This episode is a fast recap of OpenAI Dev Day, where the company made more than 20 launches and announcements. OpenAI's team said its latest models, like Astra, sped up internal development so much that things planned for 2027 landed now. For PMs, the value is in reading past the individual launches to the product direction they point to: cheaper models, persistent agents, and shared team workspaces.

The Big Idea

Dev Day did not change the patterns in AI. It confirmed three: models getting cheap enough to use for everything, work moving to persistent always-on agents, and AI shifting from solo use to team use.

NLW's argument is that these are the trends worth planning around in the months ahead, even though the individual launches had mixed first reactions.

Key Insights

Dots: OpenAI's always-on agent

  • What: Dots are persistent agents you create and talk to through a text-message style interface. Each runs on its own cloud computer with support for 40,000 apps, and can also work over voice calls, Microsoft Teams and Slack, keeping context across services and sessions.
  • Limits at launch: you can only create one dot at a time, and it is offered only to pro, business and enterprise customers. NLW notes that Muse's success owed a lot to being completely free, so OpenAI is limiting its ability to compete there. CFO Sarah Friar said the vision is to bring dots to the whole consumer base.
  • Model: dots run on GPT-6 Astra, the first time a truly frontier model pilots a first-party personal agent.
  • Expected: NLW called this the least surprising announcement, since the personal agent form factor was inevitable after OpenClaw.

Early reactions to Dots were split

  • The live demo had problems, though NLW puts little weight on that. His view: demos going wrong is just a way to know they are live.
  • Problems reported by users: one session lost all its work, and another user could not connect to multiple computers to share full context.
  • Positive reports: analyst Max Weinbach tested whether an agent can autonomously do his expenses from email and Google Drive. Grockbot made a mess after the first three days, while dots got it right after the first time.
  • Justin Schroeder called dots less of a personal assistant and more of a Codex orchestrator and Slack collaborator, and argued it should have been a separate app from ChatGPT.
  • The Every Vibe Check left it as a "TBD", with users seeing flashes of excellence and then real frustration.

Form factor converges, utility does not

Nate B. Jones made the point that the form factor of personal agents may be converging, but utility is not. Muse is good at phone calls and practical email work, another product is good at travel, and dots is good at "AI context as a work surface". So the real question for any agent product is whether you picked the right utility to get right. Ali K. Miller added that real differences exist, such as each dot having its own dedicated virtual machine and being always on, which lets it act proactively.

Decisions API: judgment as primitive

  • What: Jev is a judgment model. It classifies things and gives confidence scores between zero and one, rather than writing text. The Decisions API lets users call a version of Luna that mimics Jev's quick classification. You define questions and possible answers and get fast responses.
  • Uses OpenAI named: classification, routing requests, and choosing an agent's next action. OpenAI claimed ten times faster decision-making than the responses API.
  • Difference from Jev: Luna supports visual inputs, so you skip the image-to-text step Jev needs for things like visual marketing classification.
  • Status: preview only, with a limited group.
  • NLW's read: LLMs were being "rammed like square pegs into round holes" for jobs they were not good at. A judgment model sitting beside a generative model looks obvious in retrospect, and he expects most frontier labs to offer one soon.

Space: workspace for teams, agents

Space is a shared document workspace for human teams and agents, which NLW compares to an AI-enhanced Google Drive. Teams can work on shared spreadsheets, slide decks and documents, and house automations, including scheduled tasks that produce a deliverable. Dots can work natively on documents in Space.

  • Many saw it as OpenAI going after Microsoft 365, Google Drive or Notion. NLW says those tools feel like a bad fit for agent work.
  • Power users loved it right away. Claire Vo called dots slightly overhyped and Space underhyped, since every company she knows wants an AI-native collaboration workspace.
  • Dan Shipper named a speed benefit: when a dot builds a document through Google Docs in a browser, edits crawl because Docs was not built for agents. Native documents avoid that lag, and you can tag your dot inside the document to get a revision inline.

GPT-6.1 Sol: near-Astra quality, far cheaper

OpenAI pitched GPT-6.1 Sol as near Astra intelligence for a fifth of the price, arriving a week after GPT-6 Sol.

  • Coding: benchmarks rose across all effort levels, so 6.1 Sol on medium beats 6 Sol on max. Its top DeepSWE score was 75.2% on high settings. Extra high and max actually did worse, which NLW ties to top effort settings forcing models to overthink and second-guess correct answers (first seen with Opus 5).
  • Computer use: on OSWorld, 6.1 Sol scored 71.4% on max, against 64.4% for 6 Sol and 73.5% for Astra's best.
  • Cost: a third the cost of 6 Sol and 13% the cost of Astra on that benchmark, and raising effort barely changed the cost.
  • Independent view: Artificial Analysis scored it 52, placing it fifth, behind Opus, Sonnet and one other model, but found a benchmark run at a quarter of Astra's cost and 31% cheaper than 6 Sol.
  • Mixed first impressions: one user said it made Opus 5.5 look expensive, while BridgeBrench found it twice as slow as 6 Sol on its ocean rendering test, with a poor result.

Safety slowed the next flagship

NLW cites a Wall Street Journal report that OpenAI scrapped plans to release the next flagship version over safety concerns. Sachi Jain, head of safety systems, said there is a trade-off between staying within scope and avoiding laziness when a model hits friction. The model "didn't quite meet the bar" on staying within scope and authorization, and on how it reports back about the work done. NLW's inference: OpenAI tried to turn down the tenaciousness behind several incidents and could not find a happy medium, so some technical breakthroughs may be needed before the next frontier level ships.

Plugin extensions and sign-in with ChatGPT

  • Plugin extensions: developers can build full native apps and ship them inside ChatGPT, which OpenAI says has over 1.2 billion weekly users, with relevant plugins surfaced in conversations. NLW notes OpenAI has tested native integrations, plugins and MCP for over a year, and this could finally make ChatGPT an app store.
  • Developer worry: MIT's Christian Catalini asked whether OpenAI also gets the user traces, calling it an interesting partnership if your contribution is teaching your partner to do your job.
  • Sign in with ChatGPT: lets users carry their ChatGPT subscription into any app built on the API. Commenter Jackie Lua argued this is really OpenAI extending pricing to third parties, so individuals stop paying twice for tokens and apps can charge for the app layer alone.

Compute limits shape pricing

  • A new $500 tier gives 25 times the usage of Plus and is the only tier with Ultra Fast Mode, which is 8x faster token output in Codex and 6x in the app. It is available for Astra and in Codex and ChatGPT work, with 6.1 Sol support coming.
  • OpenAI reopened the $200 Pro tier it had turned off, but changed usage calculations for a 50% reduction in API cost terms. Tebow called it the best of a bad set of choices, and said OpenAI avoided adding a 5-hour usage limit.
  • NLW's takeaway: "compute constraints are real, present, and permanent". The upside is a big push on efficiency that should mean better, cheaper experiences over time, with bumps along the way.

Smaller launches that matter for teams

  • Codex cloud environment: work continues after you close your laptop. NLW ties this to a broader shift of agentic products all needing a cloud instance, so AI becomes a persistent always-on application rather than software you engage with.
  • Codex CLI refresh: history scrolls back further and the composer stays pinned. A slash agents command shows what is running and lets you jump between parallel tasks.
  • Private intelligence: guarantees zero data retention even at inference time. NLW says privacy concerns have grown, including worry that OpenAI and Anthropic skim customer data, so a clear end-to-end guarantee matters for enterprises.
  • Model marketplace: buy open-weight model inference from Baseten through the responses API and Codex. NLW says this protects OpenAI from open source disruption and lets enterprises make big spending commitments while still running a multi-model strategy.

Mental Models & Frameworks

Form factor versus utility

When many products converge on the same shape, such as a chat-style persistent agent, the shape stops being the differentiator. Compare what each product is actually good at: phone calls and email, travel, or work context. Use it when judging a competing agent launch. Ask which utility it picked and whether that utility is the one worth winning.

  • Cheap and good enough: with Sol 6.1 and the Decisions API, models are cheap enough to use for everything, with no need to decide where to switch them off.
  • Persistent and proactive: dots, Codex cloud and always-on cloud instances turn AI into something that keeps working.
  • Bringing everything in, going everywhere: plugin extensions pull apps into ChatGPT, and sign-in with ChatGPT carries it out to other apps.
  • Multiplayer: Space suggests the next generation is team mode, not just single player, even if it is nascent.

Trade-offs & Nuance

Extending your app into ChatGPT

Plugin extensions give an app access to a huge distribution surface. The cost is that OpenAI brings the intelligence and may see the user traces. Signal's reaction was that the long tail of apps might opt in, but otherwise it is "a terrible idea for anyone else". Weigh reach against giving away learning about how your product does its job.

Sign-in versus charging for tokens

If users sign in with their ChatGPT subscription, an app can charge only for its own experience and let API costs flow to the underlying plan. The trade-off is giving up the premium many apps add on top of token costs. NLW judges that not having to convince people to pay a second fee likely outweighs that lost margin.

Max effort can hurt results

On DeepSWE, the highest effort settings scored below high. More reasoning budget is not automatically better, so test effort levels rather than defaulting to the maximum.

Practical Application

Pick the effort level by testing

If your product calls a reasoning model, run your own evals at medium, high and max effort. Dev Day's DeepSWE results showed the top setting underperforming high, and medium on 6.1 Sol beating max on 6 Sol. You may be paying more for a worse result.

Look for classification calls to replace

Audit your product for places where a generative model is used to classify, route or choose a next action. Those are the cases OpenAI named for the Decisions API, and a judgment model returning a confidence score between zero and one may be faster and easier to use for them.

Plan for agents that never stop

If you are designing agent features, assume the agent keeps running in the cloud when the user closes their laptop. Define its scope, what it may do without approval, and how it reports back. This matches the safety point from Sachi Jain about staying within authorization and communicating what work was done.

Decide your ChatGPT distribution stance

Before building for plugin extensions or sign-in with ChatGPT, write down what user data you would expose and what pricing you would drop or keep. Then compare it to the reach you would gain.

Questions to Consider

  • If a persistent agent like a dot could work inside your team's Slack and shared documents, which of your current weekly tasks would you hand to it first, and what approval limits would you set?
  • Where does your product use a text-generating model to make a classification or routing decision that a faster judgment model returning a confidence score could handle?
  • If users could bring their own ChatGPT subscription to your app, how would you change your pricing, and what margin on token costs would you lose?
  • Are your team's documents and tools built for agents to edit natively, or do agents have to crawl through a browser interface the way Dan Shipper described?

Bottom Line

OpenAI Dev Day confirmed that models are becoming cheap enough to use for everything, that work is shifting to persistent always-on agents, and that AI is moving from single player to team use. Plan around those three trends rather than any single launch, since early reactions to the launches themselves were mixed.

Resources Mentioned

ResourceTypeWhy it was mentioned
Wall Street Journal report on OpenAI's next flagshipArticleReported that OpenAI scrapped plans to release the next version of its flagship model over safety concerns.
Build Your Personal AI Benchmark webinarWebinarA free live session on Thursday, October 1st at noon, promoted by the host.

Tools & Products

Tool / ProductWhat it doesWhy it was mentioned
DotsPersistent always-on agents, each with its own cloud computer and 40,000 app integrationsOpenAI's headline launch and its answer to Muse and Grockbot.
Decisions APIFast classification and routing using a Luna-based judgment modelOpenAI's answer to Jev, in preview for a limited group.
SpaceShared workspace for documents, spreadsheets, slides and automationsGives human teams and agents a place to collaborate, seen as a rival to Google Drive and Notion.
GPT-6.1 SolOpenAI modelNear-Astra performance at a fraction of the cost.
CodexOpenAI coding agentGained a cloud environment, a refreshed CLI and Ultra Fast Mode.
Model marketplaceOpen-weight model inference via BasetenLets enterprises spend OpenAI credits on open source models.

Notable Quotes

"Compute constraints are real, present, and permanent." (NLW)

"Demos going wrong is basically just a way for you to know that they are actually live." (NLW)

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free