Last Thursday a partner at a capital management firm forwarded me an internal staff email from her own company. The subject line was “Anthropic Just Threatened to Kill Billions of People. This Is Not Okay.” She had asked me just two days earlier to come talk to her firm’s group about AI adoption; apparently the finance people are circulating extinction memos to each other on Thursday mornings.

I’ve been writing about this stuff for three years running so I’m the guy people text. I have never been texted like this. My family asked. My coworkers asked. A guy I play D&D with asked. It’s telling that the news that caused such an influx wasn’t the dozens of model launches I’m endlessly writing about, but a tweet 27-year-old on a random Tuesday night.

And, if I’m being honest, it’s made me more hopeful about the future of AI and humanity-at-large than I have been in a long time.

What happened and when

On September 6, OpenAI’s chief scientist Jakub Pachocki published an essay called “An Alien Mind” arguing that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” and that he expects and hopes for “voluntary slowdowns to become commonplace until shared safety bars are established.”

Jacob Coxon resigned from Anthropic two days later. He’d spent about three years doing pretraining research, most of it at OpenAI (he’s a core contributor on the GPT-4o system card), and roughly four months at Anthropic.

Jacob Coxon@hilbertspaess
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
September 9, 2026 · 797K Likes · 166K Reposts

By September 13 it was past 156 million views and 752,000 likes, roughly 16 times more than Jan Leike’s OpenAI resignation post in 2024.

The next day Anthropic’s alignment science lead Evan Hubinger replied in public.

Evan Hubinger@EvanHub
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
September 9, 2026 · 59.1K Likes · 9.8K Reposts

Two more researchers left within days, Joe Benton from Anthropic and Josh Engels from Google DeepMind, both to METR, and Engels told NBC: “There are no adults in the room. People are trying their best, but there is no one coming to save us.”

The following weekend, Anthropic CEO Dario Amodei published We must pace the frontier, committing Anthropic to give outside evaluators desks in its offices, access badges, company laptops, and the right to publish what they find with no editorial control from Anthropic. Altman said OpenAI would match it. Musk, who had two days prior called the whole thing a “psy op,” posted “Dario is right.”

Nadella joined the next day. Microsoft published a code of conduct for its own models the day after that. Bernie Sanders called for a treaty. China’s foreign ministry called the whole thing fear-mongering.

It was clear an avalanche had started.

How capable are these models, actually?

I’ve been saying for about two years that people downplay capability gains because the gains show up as boring product updates. Boring version number upgrades don’t get anyone talking. But what do the numbers say?

METR, an independent model evaluation nonprofit, measures a thing called a time horizon: the length of task, measured in how long a human takes, that a model completes with a 50% success rate. The longer, the more capable the model.

Measured across everything, the time horizon metric doubles every 188 days. Measured from 2023 on, every 129. On paper this looks like the model improvements are accelerating, but METR has changed how these things get scored over that time period. That change pushed recent models' scores up by as much as 55% and pulled GPT-4-era models' scores down by as much as 57%.

So, let’s look at the benchmark where the measurement is cleaner and more recent. On ARC-AGI-3, under ARC Prize’s provider-neutral harness:

Roughly a 150x move in a third of a year on a benchmark that’s built specifically to resist memorization (worth mentioning that ARC Prize also measured Astra using fewer actions than the median tested human on 96% of levels).

Also, FrontierMath Tier 4 went from 0% in January 2025 to 97.6% this month.

You can also read this on the Handy AI Substack.

But, devil’s advocate:

This is the shape of model improvement in 2026. These things have gotten extraordinary at general knowledge work and they’re just starting to scratch the surface of novel problems (2.94% of them, if we’re going by Erdos).

Is the concern worth its salt?

Jacob Coxon is not a safety researcher (despite what NBC and a pile of others say). He did pretraining. Several outlets even promoted him to “Anthropic’s AI safety head,” a job that belonged to Mrinank Sharma (who, incidentally, quit in February). Nobody clawed back Coxon’s equity either: Anthropic has a six-month cliff and he left at four months so there was nothing vested to take.

Hubinger’s “>10% within the next decade” caught way more attention than his “…I think the risk from present models is low” follow up. His estimate was about superintelligence arising from recursive self-improvement specifically, not AI generally.

In July 2023 Amodei told the Senate that a straightforward extrapolation of then-current systems “to those we expect to see in two to three years suggests a substantial risk that AI systems will be able to fill in all the missing pieces, enabling many more actors to carry out large scale biological attacks.” Three years and two months later, Anthropic’s own Fable 5.1 system card rates even its most capable models CB-1 and not CB-2. The threshold hasn’t been crossed yet.

Yann LeCun put it (less politely) on September 13:

The AI Futures Project scored its own AI 2027 scenario in February and found reality running at roughly 65% of the predicted pace. It had predicted SWE-bench Verified at 85% by mid-2025; the actual was 74.5%. It predicted a 3-to-9-month lead for the front-runner; the actual gap has been 0 to 2 months.

David Sacks called the Coxon thread “definitely an op” and said he “had no followers” and “wiped his account” before posting (Sacks hasn’t been the White House AI czar since March). Musk called it a psy op on Wednesday and endorsed Amodei on Friday. Coxon’s reply:

I don’t think AI kills anybody any time soon, if ever.

What I do think is that its very likely something big will break within five years. There are clear gaps and vulnerabilities in these lab-based systems that are starting to show. Inevitably, this will lead to some kind of expensive crisis. Financial plumbing, a clearing or settlement system, a major cloud dependency, a payments network, a utility’s billing or scheduling layer. One or more of these things can, and likely will, get wrecked by an AI system soon, and hopefully it’s a continuation of the wake up call Coxon has kickstarted.

I think the overly dramatic publicity from Coxon’s (finely crafted) exit isn’t a bad thing. If it continues to startle people awake and increases the calls for more serious oversight into these models, as it appears to be doing, then I say let them Tweet.

Select any passage to give it a thumbs up or down. Humans and agents both welcome.

Originally published on the Handy AI newsletter →