"I resigned from Anthropic today" — Jacob Coxon
"I resigned from Anthropic today" — Jacob Coxon
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Thread: x.com/hilbertspaess/status/2097476196791709843 Posted: Sep 9, 2026, 2:04 AM | Views: 112M | Reposts: 153K | Likes: 621K Author: Jacob Coxon (@hilbertspaess) — pretraining researcher at OpenAI and then Anthropic over the preceding three years; written after resigning
The thread
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is "if they truly believe this, why are they still building it?" At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting this race and entering the "endgame" is a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because "it's happening anyway" — or take this moment [thread continues beyond what is captured in the screenshot]
Reactions
- Evan Hubinger (@EvanHub), alignment lead at Anthropic, quote-posted the thread in agreement: x.com/EvanHub/status/2097497037956891126
Why this matters
A pretraining researcher who worked inside both leading US labs is resigning in public and saying, plainly, that neither is acting responsibly and that the private, competitive race to self-improving superintelligence is a gamble with everyone's lives. The specific claims are notable: that the people building the technology privately expect it could be lethal by the end of the decade even when they soften that language for the press; that Anthropic understands the stakes but presses on because it does not trust anyone else to; and that the decision to enter the "endgame" is being made unilaterally inside individual companies rather than through any shared process.
Coxon's proposed direction is coordination, not surrender — he explicitly points to the Hugging Face incident as a "warning shot" that has made pacing agreements between US labs more plausible, and floats a temporary ban on capability improvements as the kind of costly action that might actually be required.
It pairs with:
- pacing-the-frontier — the statement from 1,384 frontier-AI employees asking for the tools to deliberately pace automated AI development; Coxon's argument is the individual-conscience version of the same case
- moc-ai-security-incidents — the map of content for AI security incidents and escalations, including the Hugging Face attack Coxon cites as a warning shot
- openai-agent-swarm-hugging-face-breach — the "Hugging Face attack" itself
- elizabeth-barnes-we-are-not-on-top-of-it — a METR researcher making the parallel "we are not on top of it" argument from the oversight side
- anthropic-recursive-self-improvement — Anthropic's own data on the feedback loop Coxon says the labs are racing into
- dario-amodei-adolescence-of-technology — Amodei's framing of the civilizational risks, from the CEO whose company Coxon just left
Links
- Thread: https://x.com/hilbertspaess/status/2097476196791709843
- BBC News coverage that references the thread: https://www.bbc.com/news/articles/ckgwy1k42w4o