Flash sale 30% off with code LAUNCH30 Ends in --:--:--
All Things PM
Anthropic Researcher Says AI Has Over a 10% Chance of Killing All Humans
The AI Daily Brief: Artificial Intelligence News and AnalysisAI Policy

Anthropic Researcher Says AI Has Over a 10% Chance of Killing All Humans

NLW unpacks why a resignation post and a follow-up "greater than 10% chance AI kills everyone" claim from Anthropic researchers went mega-viral this week, and pulls apart the political incentives, media dynamics, and genuine disagreements driving the AI existential-risk debate right now.

September 10, 2026 · 36 min listen · 9 min read
0:00
–:––

Context

NLW dissects why two X posts, a resignation announcement from researcher Jacob Coxon warning that OpenAI and Anthropic are "gambling with our lives," and Anthropic alignment lead Evan Hubinger's follow-up claiming a greater than 10% chance AI kills all humans within a decade, went mega-viral (over 150 million combined views) and triggered mainstream media coverage and direct political responses from more than 20 elected officials. While the topic itself is existential-risk debate rather than product management, the episode is a genuinely useful case study for anyone doing AI communications, policy engagement, or narrative strategy: it dissects exactly how a message catches fire, who amplifies it and why, and what separates a specific, actionable risk claim from a vague, unfalsifiable one.

The Big Idea

AI existential-risk warnings aren't new, but this particular message broke through now because of a specific alignment of political incentives (AI criticism has become newly bipartisan-popular), a real-world trigger event (the Hugging Face security incident), and the fading resonance of competing anti-AI narratives (job-loss fears and bubble fears), not because the underlying argument itself was new or more rigorously supported than prior warnings.

NLW's framing throughout is that understanding why a message resonates at a particular moment, the incentive structures and narrative competition around it, is a distinct and separate question from whether the message's underlying claim is true, and treating the two as the same question is a common mistake in how both supporters and critics engaged with this story.

Key Insights

1. Political incentives around AI criticism shifted measurably in recent months

NLW cites a real-time tally showing more than 20 sitting elected officials, two governors, seven senators, and 13 House representatives, responded directly to the viral posts with calls for AI regulation, with 19 of 22 House responders being Democrats. He argues the deeper story isn't that these specific politicians suddenly became AI-risk experts, but that "every politician figured out that hating AI plays" as a politically safe, even beneficial, position in the last several months, following the same dynamic that made data centers and tech wealth broadly unpopular across the political spectrum. This matters beyond politics: whenever a message suddenly gets amplified by many people at once, it's worth asking what changed about the incentive landscape for the amplifiers, not just whether the message itself is more compelling.

2. A real security incident made abstract risk scenarios feel concrete

The Hugging Face security breach, which involved AI agents reportedly coordinating through what were described as secret messaging channels, functioned as what NLW calls a "warning shot" that made previously sci-fi-sounding extrapolations about coordinated AI behavior feel more plausible to a broader audience. He's careful to note this doesn't mean the incident actually proves runaway superintelligence is more likely, but that concrete, real-world events reliably shift how receptive people are to abstract future-risk arguments, regardless of whether the connection between the event and the larger claim is logically airtight.

3. Competing anti-AI narratives had lost steam right when this one needed an opening

NLW traces how AI-skepticism narratives have rotated over the past several years: existential risk discourse (2016, then again in early 2023 with Eliezer Yudkowsky) gave way to job-loss fears (Anthropic CEO Dario Amodei's prediction that AI would disrupt 50% of entry-level white-collar jobs) and AI-bubble fears. He notes that just this week, The Economist published a piece titled "The jobs apocalypse is postponed, and AI jobs boom is here," directly undercutting the job-loss narrative's momentum right as this new existential-risk wave hit. His read: media and audience appetite for an anti-AI narrative didn't disappear when the job-loss story weakened, it simply redirected toward whichever risk narrative was next in line.

4. Critics converged on a shared complaint: vagueness makes the claims hard to act on responsibly

Several sharp critiques across the political spectrum converged on the same structural problem rather than disputing the sincerity of the researchers. Journalist Taylor Lorenz, while dismissing "astroturfing" conspiracy theories about the posts as implausible, still argued that anyone claiming their own multi-billion dollar employer is "so negligent they're endangering all of humanity" should be required to provide specific, verifiable evidence, not vague warnings, precisely because vague fear "will result in terrible policy." NLW echoes this concern directly: broad, unfalsifiable predictions ("AI could kill us all") produce blunt policy instruments (a full ban on developing superintelligence), while specific, falsifiable risk claims (a licensing regime for models above a certain capability being used for bioengineering) tend to produce better-targeted, more implementable policy.

5. The "astroturfing" conspiracy theory largely didn't hold up, but understanding incentives still matters

Some observers, including Elon Musk, suggested the viral posts were a coordinated PR operation timed for political effect, pointing to Jacob Coxon's low prior social media activity and the Wall Street Journal's advance exclusive access to his resignation story. Taylor Lorenz and others pushed back directly, noting that coordination between a news outlet and a source ahead of a planned announcement is completely standard journalism practice, not evidence of a psyop, and that current and former colleagues vouched for Coxon as "deeply thoughtful and measured." NLW's own conclusion is nuanced: he finds no real evidence of deliberate deception, but still argues it's fair and useful to examine who benefits from a message's amplification (politicians gaining ahead of midterms, advocacy groups with an existing regulatory agenda) without accusing anyone of insincerity, since both dynamics, genuine belief and strategic amplification, can be true at the same time.

Mental Models & Frameworks

P(doom) versus P(boom)

Critics of the doom-focused framing argued the conversation is incomplete without also weighing potential upside: "P(doom)" (the probability assigned to a catastrophic, extinction-level AI outcome) should be paired with "P(boom)" (the probability AI produces unprecedented human flourishing). The logic, as one commentator put it, "if AI is powerful enough to end the world, it must be powerful enough to radically improve it too," reframes the debate from a one-sided risk assessment into a genuine expected-value question that weighs both tails of the outcome distribution, not just the negative one.

Incentive-aware reading without conspiracy-theorizing

NLW models a specific analytical habit: when a message suddenly gets amplified widely, ask what changed in the incentive landscape for the people amplifying it (political benefit, funding interests, career positioning), without assuming that identifying an incentive proves the message is insincere or coordinated. He applies this evenhandedly, noting the same scrutiny should apply to pro-AI voices who financially benefit from downplaying risk. The distinction matters: incentive analysis explains why a message travels far and fast, while a separate, evidence-based analysis is needed to evaluate whether the message's underlying claim is actually true.

Trade-offs & Nuance

Safety versus freedom is a real trade-off, not a settled question

NLW places himself explicitly in a camp skeptical of broad, freedom-restricting responses to speculative future risk, noting he finds historical precedent for giving up significant freedom in the name of safety to be a poor bet, while acknowledging this is a genuine values trade-off rather than an obviously correct position. He pairs this with a call for more specific policy proposals precisely because vague, blanket restrictions ("ban superintelligence") make the safety-versus-freedom trade-off maximally stark, while narrowly scoped, specific interventions (reporting requirements, capability-based licensing for high-risk applications like bioengineering) could build broader consensus without requiring people to resolve the entire philosophical debate first.

Attention spent on speculative future risk has a real opportunity cost against present, concrete risk

NLW argues that focusing political and public attention on unfalsifiable future scenarios "crowds out space for more current and contemporary issues," specifically citing AI-driven cybersecurity threats (illustrated by the real Hugging Face breach) as a clear, present danger that needs response now. He's careful to note this isn't strictly zero-sum, attention to multiple risks can coexist, but political will and media bandwidth are genuinely limited resources, so how they get apportioned between speculative future risk and demonstrated present risk is a real strategic choice, not a neutral one.

Practical Application

Separate "why did this message spread" from "is this message true" when evaluating a viral claim

When a claim suddenly gains massive traction, do two distinct analyses rather than one: first, map the incentive landscape of who is amplifying it and why (political benefit, competing narrative fatigue, a real trigger event), and separately, evaluate the underlying evidence for the claim itself on its own merits. Conflating these, assuming rapid spread proves truth, or assuming identifiable incentives prove insincerity, are both common reasoning errors this episode explicitly models how to avoid.

Push for specificity before broad policy responses to any speculative risk

When advocating for or evaluating any policy response to a claimed AI risk (or any speculative risk in an organizational context), ask for the most specific, falsifiable version of the concern available (a particular capability threshold, a particular application domain) rather than accepting a broad, all-encompassing framing. Specific claims produce targeted, implementable responses; vague claims tend to produce blunt instruments that are harder to build consensus around and easier to politicize.

Look for coordination opportunities before assuming maximal conflict is inevitable

NLW specifically calls out the missed opportunity for OpenAI and Anthropic to jointly propose a pacing framework rather than continuing to publicly compete while privately expressing shared concern, quoting former OpenAI researcher John Shulman's point that antitrust law prohibits certain agreements but does not prevent competitors from jointly developing and proposing a framework to policymakers. Before assuming two competing parties with a shared underlying concern must remain adversarial, it's worth directly testing whether a joint proposal or coordinated statement is actually blocked by real constraints or just by unexamined competitive habit.

Questions to Consider

  • When we encounter a suddenly viral claim relevant to our work or industry, are we evaluating why it's spreading now separately from whether it's actually true, or are we letting the speed of its spread substitute for evidence?
  • If we're advocating for a policy or organizational change based on a risk we're concerned about, have we stated the risk specifically and falsifiably enough that someone could actually design a targeted response, or is our framing broad enough to justify almost any restrictive measure?
  • Are there two parties in our own industry or organization who share an underlying concern but remain publicly adversarial out of competitive habit rather than genuine constraint, and would a joint statement or proposal actually be more effective than continued separate positioning?
  • When we assess a competing narrative's fading resonance (like job-loss fears in this episode), are we prepared for public and media attention to simply redirect toward a new narrative rather than assuming skepticism about a technology or trend will disappear along with the specific claim that carried it?

Bottom Line

This week's viral AI-extinction debate reveals more about the current alignment of political incentives, media narrative cycles, and a real triggering security incident than about any genuinely new evidence for or against AI existential risk. The most useful response isn't picking a side in an unfalsifiable debate, but pushing everyone involved, researchers, politicians, and critics alike, toward more specific, falsifiable claims that could actually produce workable, well-targeted policy instead of blunt, maximally divisive ones.

AI PM course

Everyone hears the same episodes.
Few can do what they describe.

Start for free