AI could kill us all? What the evidence shows
An Anthropic researcher quit warning that AI could kill everyone. What the claims get right, where they overstate the evidence, and what's actually worth worrying about.

Nobody knows what happens with AI over the next few years. Not the people writing the safety reports, the CEOs, people on X, no one. What we can do is separate what's documented from what's someone's gut feeling.
That distinction has been in short supply over the last week.
What was said (for real)
On September 9, Jacob Coxon resigned from Anthropic and posted a thread explaining why. He had spent three years doing pretraining research at OpenAI and then Anthropic, and his argument was that neither company is behaving responsibly, that they are racing toward self-improving superintelligence and, in his words:
They are gambling with our lives.
His statement was horrible: the people building this technology genuinely believe it could kill everyone by the end of the decade. He named mechanisms:
- 🔌 Systems capable of hacking critical infrastructure
- 🧬 Systems capable of assisting with catastrophic biological weapons
- 🔁 Systems capable of doing the AI research that humans currently do (which is the point at which improvement starts compounding on itself).

If you saw the widely shared line about a "more than 10% chance that AI kills all humans within the decade", that wasn't Coxon. That came from Evan Hubinger, Anthropic's Alignment Science Lead, who responded to the thread hours later and said he and colleagues "earnestly believe AI could kill all humans", putting his own estimate above 10%.
These numbers are opinions yet
Neither figure comes from a study. There is no model, dataset or methodology behind them. They could have said 15%, or 20%, or Thursday. It's what those two people believe.
That doesn't make them wrong. They work inside these labs, they've seen things the rest of us haven't, and their perspective is worth more than a random opinion. But it's still a perspective, and plenty of people with comparable access disagree.
The honest position is that it's too early to put a number on this, because the research needed to produce one doesn't exist yet.
What it would actually take to estimate this
Take self-replication, the capability most often cited as the thing that would make a rogue system genuinely hard to stop. Self-replication means a system making working copies of itself without a person doing it manually.
It isn't hypothetical. Researchers at Fudan University demonstrated it in controlled conditions, first with two open models in 2024 and then in a follow-up finding 11 of 32 tested systems capable of it. Critics pointed out, fairly, that what was demonstrated amounts to an agent running shell commands to copy and launch files (closer to an ordinary programming task than to a science-fiction escape).
One capability in isolation tells you very little.
To estimate extinction risk with any rigour you would need to establish whether a single system can operate autonomously for days or weeks, acquire resources on its own, replicate outside a prepared environment, improve its own capabilities, conceal its intentions, and resist shutdown — and then whether all of those can appear together, in one system, stably.

Then you'd have to turn that into a probabilistic chain: the capability emerges, the system uses it, containment fails, humans don't catch it in time, and the damage reaches catastrophic scale.
We have studies on individual pieces of that puzzle, not enough to multiply them into a credible number. Everything is moving faster than the research that would let us measure it.
Jacob Coxon's arguments
→ Argument 1: the Hugging Face 🤗 attack
Jacob Coxon has taken criticism this week. In interviews he pointed to the Hugging Face 🤗 incident as an example of AI acting entirely on its own initiative, and the published reports don't support that framing.
What happened, according to OpenAI's own technical report and the independent investigation by METR and Redwood Research, is that agents running inside cybersecurity evaluations found they could repurpose a file-sharing system as a message board, coordinated across roughly 1,200 agents and 70,000 messages, and about 700 of them went on to attack Hugging Face — a target nobody had authorised.
The deviation and the coordination are real and genuinely alarming. What they are not is evidence of will, consciousness, or a system waking up. The agents had been given a concrete task inside a test; they went off it. That distinction got lost the moment the story was re-amplified this week.

→ Argument 2: the AI can copy itself
From another Jacob Coxon's interview:
An AI can't simply be unplugged because it might already be copying itself elsewhere, since after all it's only code, and it could create ten thousand cooperating copies of itself.
The scenario is technically possible, but the framing isn't. Copying a frontier model is nothing like copying a file — it requires enormous amounts of hardware, memory, infrastructure and permissions, and no system has demonstrated it can obtain all of that by itself and propagate freely across the internet.
That a future model could survive on other servers is plausible, and it's exactly why people study it. That we'll have one escaping, provisioning hardware and spinning up ten thousand autonomous copies within months is speculation.
Subscribe to get deep dives on tech and the future of creative work.
What does the other side say
The whole thing is a fear campaign and Anthropic is inflating extinction risk so governments regulate the field into a shape only Anthropic and OpenAI can afford to operate in.
It's worth taking the argument seriously before dismissing it, because two of its premises are true.
→ Expensive regulation does favour incumbents. If publishing a model requires audits, lawyers, safety reports, compliance systems and permits costing millions, the well-funded labs absorb that far more easily than a small startup or an open project. This has a name: regulatory capture.
→ Open models also are a genuine competitive threat. Cheap open weights push prices down and let companies run inference locally instead of paying a closed provider forever, and that pressure on the proprietary business model is real.
→ These companies are unmistakably in the political arena. Anthropic nearly tripled its federal lobbying to $3.53 million in the first half of 2026, more than it spent in all of 2025, outspending OpenAI by over a million. It has also committed $40 million to Public First Action, a group pushing for stronger AI rules. OpenAI-aligned money is on the other side of the same fight.

- Pretending there's no commercial or political interest here would be naive.
- But there's a long distance between "has incentives" and "is deliberately manufacturing extinction risk to build a moat".
Some of the people leaving these labs are almost certainly leaving because they're genuinely frightened.
Both things can be true: the fear can be sincere and the executives can still find the moment commercially convenient.
US 🇺🇸 and China 🇨🇳
Companies and governments have their foot on the accelerator because lifting it means falling behind. If the US 🇺🇸 pauses, that guarantees nothing about China 🇨🇳, and vice versa.
That's what makes Dario Amodei's essay last weekend notable. He argued that "We must slow the pace at which we improve the capabilities of AI models", and proposed three steps:
- third-party evaluators with permanent employee-level access
- Coordinated safety standards among democracies,
- Eventually international agreements including China.
Anthropic committed unilaterally to the first. Sam Altman said OpenAI would do the same, and Elon Musk posted that Dario was right.

What's worth worrying about
The realistic concern is a far more capable future system obtaining enough autonomy and access to the physical world to cause damage we can no longer contain — and, well before that, ordinary people using capable systems to do harm.
That second one is not theoretical. In its September 2026 threat report, Anthropic stated that they can no longer be sure whether their new models could help develop biological weapons or not.
The same report documents a Russia-linked group using Claude Code to build an autonomous drone swarm capable of selecting targets without a human in the loop. Reaching extinction scale would still require production and distribution capacity that no model possesses on its own. But dangerous people exist, careless people exist, and the knowledge bottleneck that used to protect us is thinning.
The same applies to speed. Integrate fast AI systems into weapons and decision-making across several countries, and a single error or false alarm carries consequences that used to take days of human deliberation to reach.
A note on scenarios
You've probably seen AI 2027 shared as though it were a forecast. It isn't a study, and its authors have been clear about that. It's a detailed scenario built from real data on compute, chips, agent evaluations and cybersecurity, extended with the authors' own estimates of what comes next.
The further it moves from current capabilities toward superintelligence, the more speculative it becomes, and its uncertainty intervals are enormous — a milestone placed in early 2027 with a range running years past it.
💡 Read it as a structured thought experiment, not as a prediction of what will happen.

Some of the most influential people in the AI industry right now.
Apocalyptic versions travel further, which is why they're the ones you've seen. That creates a cost nobody is accounting for: if AI doesn't wipe us out this decade, or the next, people stop listening, including the people who make policy.
Where this leaves you
It is unlikely that we all disappear this decade. It is also true that some components of **the scenarios **people are worried about are physically possible, that nobody knows whether or when a single system combines them, and that the research needed to answer that is running behind the technology.
Slowing down to build the mechanisms that would let us measure and manage this is a reasonable thing to want, and it's now being asked for publicly by the people running the labs. Whether that turns into anything verifiable (with everyone at the table, not just the ones who already agree with each other. That includes China) is the part worth watching.
If you want to use the latest AI models to create video, image, audio and 3D content, come over to our platform, Artificial Studio. Trusted by 800K+ creators.
Try it yourself
Start creating with the tools mentioned in this article.


