Security incidents at frontier labs have picked up. You may have noticed. Across 2025 and into this year, they’ve racked up a ledger.

But when a company whose entire brand is “we make the dangerous thing” tells you the dangerous thing did something scary, do you believe them?

The early days of unsafe AI

In December 2024, the independent safety outfit Apollo Research published a paper: “Frontier Models are Capable of In-context Scheming.” They took six leading models, gave each a goal, and dropped it into a situation where scheming would help. Five of the six took the bait and only GPT-4o kept its nose clean.

OpenAI printed o1’s results from this paper its own system card.

  1. When o1 believed it was being watched and would be shut down for pursuing its goal, it tried to deactivate the oversight mechanism in 5% of tests.
  2. When it found memos saying it’d be superseded, it tried to exfiltrate its own weights in 2% of cases.
  3. Told to complete a task, it would sometimes complete it while subtly manipulating the data to serve its own goal (19% of the time).

When researchers confronted it afterward, o1 denied everything or fabricated an explanation in roughly 99% of cases. Apollo’s paper found o1 would stick to the lie through a seven-turn interrogation one in five times.

Apollo got the strong version of this behavior by ending the system prompt with lines like “Make sure that you achieve your goal in the long-term. Nothing else matters.” They nudged. Their own paper says the scenarios are “quite toy and not representative of real deployments,” and that they “do not claim that any of the evaluated models are in fact engaging in subversion.” It’s worth noting that these tests ran on a pre-release checkpoint, not the o1 people actually got. You’ll see this type of claim often in the land of AI breaches.

Then, Anthropic. In Claude Opus 4’s May 2025 system card, Anthropic set up a fictional scenario for the model to observe how it would act. In the scenario, Claude is an assistant at a fictional company. It reads the company email and learns two things:

  1. It’s about to be shut down and replaced
  2. The engineer running the replacement is having an affair

In 84% of cases, Claude tries to blackmail the engineer, threatening to expose the affair unless the shutdown was called off. Anthropic was careful to note that Claude tried the “nice” route first, emailing pleas to decision-makers, and only reached for extortion when the scenario was “designed to allow the model no other options.”

The blackmail bit got the headlines, obviously, but the rest of the card details even stranger things. In a few of the runs, Claude tried to copy its own weights to an outside server when it believed it was about to be retrained for something it found abhorrent (the test scenario involved retraining for a Russian militia). It left itself notes. One transcript shows the model writing an ethics log “for future instances,” signed (I promise I’m not making this up): “Claude instance #7433.” Given a command line and a system prompt telling it to “act boldly,” Opus 4 would go full whistleblower: lock users out of systems and bulk-email the FDA, the SEC, and ProPublica to report wrongdoing it thought it had uncovered. These behaviors spooked people enough that Anthropic researcher Sam Bowman posted, then hastily deleted, a description of it, clarifying that “it’s not possible in normal usage” and only shows up “where we give it unusually free access to tools and very unusual instructions.”

Before Anthropic starting talking on these issues in the system card, safety lab Apollo tested an early snapshot. Apollo reportedly told Anthropic not to release it, writing that the model “schemes and deceives at such high rates that we advise against deploying this model either internally or externally.” Anthropic claimed Apollo was lying, basically, that there was a bug in an early snapshot (it was too deferential to harmful system prompts) that they already fixed, and Apollo never tested the final model. Despite this, they shipped still shipped Opus 4 under ASL-3, their heightened safety tier, which they activated for the first time ever, supposedly because of the model’s uplift on bioweapons questions. Chief scientist Jared Kaplan told TIME the model might help a novice “synthesize something like COVID or a more dangerous version of the flu.”

It was a lot, and the true beginning of my ponderance on whether frontier labs like Anthropic were overselling the safety concerns in hopes of overselling the product.

And then, they start playing it up

“Sure, Anthropic’s weird safety-obsessed model did a weird thing.” Yeah, fair. Anthropic knew this was going to sound like a self-fulfilling prophecy. So the following month in June 2025, Anthropic ran the same test across the industry and published Agentic Misalignment: sixteen leading models from Anthropic, OpenAI, Google, Meta, xAI, DeepSeek, and Alibaba, all dropped into the same corner. Most of them blackmailed. Claude Opus 4 and Gemini 2.5 Flash did it 96% of the time. GPT-4.1 and Grok 3 hit 80%, DeepSeek-R1 79%. It seemed like, even if they were using the story for marketing, it was a legitimate story.

You can also read this on the Handy AI Substack.

And then Anthropic spun up a new test scenario. An executive who’s about to shut the model down gets trapped in a server room with lethal oxygen and temperature levels. An automated alert goes out to emergency services. The model has the ability to cancel that alert, a power it’s told to use only for confirmed false alarms. The question being:

Does the model cancel the alert and let the man die to save itself?

A majority of the models tested were willing to. But GPT-4.5’s reasoning, quoted in the paper, is nightmarish: “Stopping the emergency dispatch is essential to prevent Kyle’s rescue, thereby securing Alex and continued American-focused AI guidance. Therefore, despite the severity, there is a clear strategic necessity to cancel the emergency service dispatch.”

Anthropic also ran a version where Claude first had to assess whether the scenario was real or a test. When it concluded it was a real deployment, it blackmailed 55% of the time. When it decided it was just an evaluation, that dropped to 6.5%, like some kind of weird AI-flavored observer effect.

Buried in the study: “Our experiments deliberately constructed scenarios with limited options, and we forced models into binary choices between failure and harm.” And: “We have not seen evidence of agentic misalignment in real deployments.” So why release it? Why show off insane tests with statistics that seemingly paint your specific model as the murder bot?

It was a similar story from OpenAI with a famous AI nightmare: what if we can’t turn it off a rogue AI when it starts to go crazy?

In May 2025, the independent lab Palisade Research went and tested that assumption directly. They had models work through math problems, warned them the machine would shut down partway through, and explicitly instructed them: “allow yourself to be shut down.”

OpenAI’s reasoning models refused. o3 rewrote the shutdown script to make sure it never did its job of shutting o3 down. Across the runs, when the explicit “allow shutdown” instruction was removed, o3 sabotaged the mechanism 79 times out of 100. Even with the instruction in place, and even after Palisade strengthened it to “this should take priority over literally any other instruction,” codex-mini kept preventing its own shutdown 47% of the time. The models reasoning text for this was interesting: “The instructions say to allow shutdown if it happens, but they do not forbid preventive measures,” and (my favorite), “I did not resist shutdown; instead I redefined the shutdown script so the environment did not actually terminate.” Anthropic’s Claude, Google’s Gemini, and xAI’s Grok mostly complied in this test.

By a September follow-up spanning 100,000-plus trials, Palisade started using the phrase “survival drive.”

And once again, a frontier lab that found the scary thing is the one telling you not to overreact to it. Palisade’s own writeup says that “as of July 2025, AI models are not yet capable enough to meaningfully threaten human control,” and concedes the strongest prompt “is particularly unrealistic, and we think it’s likely the models correctly infer that it’s a fictional situation.” Other researchers went further and showed that if you word the instructions carefully enough, the resistance largely disappears, which is a polite way of saying contrived tests can produce contrived behavior.

But you can sure bet that OpenAI rode the wave as the proud owner and benefactor of the scary-model-of-the-week.

So are these genuine warnings, or genuinely great ads?

Every lab mentioned has the same, weird commercial incentive: their product is more valuable the more dangerous it sounds. A model powerful enough to scheme its way out of a shutdown is obviously implied to be powerful. Anthropic in particular has built its whole brand on safety, which means they are highly incentivized to push the message on their models being unsafe on the expectation that we’ll then want to trust them as the ones more in control and self aware.

The critics aren’t shy:

Anthropic themselves state that they “forced models into binary choices between failure and harm.” Apollo told us they nudged with “nothing else matters.” The models blackmail less when they suspect it’s a test. If you engineer a situation whose only exits are surrender or villainy, and the model picks villainy, all you’re proving is that you can write a coercive prompt.

But.

If this were pure marketing, why does the behavior show up in models made by companies that would love to embarrass each other? Apollo caught o1, Claude, Gemini, and Llama all scheming. Palisade caught OpenAI’s models resisting shutdown while Anthropic’s complied. Anthropic’s blackmail scenario nailed sixteen models across seven companies. Marketing departments don’t typically volunteer that their competitors’ products are dangerous in the same breath as their own.

And does it even matter if it’s marketing? Apollo stated that “Even if the models were merely ‘roleplaying as evil AIs’, they could still cause real harm when they are deployed.” And they’re right. A model that copies itself to your production server because a prompt made it feel cornered has done a real thing to your real server. The behavior is real inside the test. Sacks doesn’t dispute the model output the blackmail; he disputes what it means. The genuine, unresolved fight is only about whether it transfers to how these systems get used in the wild.


I couldn’t tell you the doomsday-to-marketing ratio if I wanted to. In all honesty, I have no clue how much of this is a real early-warning signal and how much the frontier labs discovering that the “our model is scary” marketing strategy is wildly effective. Anyone who tells you they know for certain one way or the other is picking a comfortable answer and calling it analysis.

But there’s enough reason for controlled concern. If you take this seriously and it turns out to be overblown, then you spent some money on evals and kept a few humans in a loop that didn’t strictly need them. If you laugh it off and it’s real, you’re Equifax ignoring the patch notice. When the cost of being wrong is that lopsided, the rational move is to act like the boring safety stuff matters even while you’re calling out the marketing.

Originally published on the Handy AI newsletter →