Jacob Coxon quits Anthropic: AI could kill us all

Jacob Coxon resigns from Anthropic: 'AI could kill us all' by the end of this decade. A colleague at the lab puts the risk above 10%.

English · Original discussion in Spanish · Published

Coxon quits Anthropic: 'AI could kill us all'

Coxon was 27 years old and held a job that half the tech industry would kill for: a researcher at Anthropic, the lab that presents itself to the world as the cautious version of OpenAI. He has left. His argument fits in one sentence and admits no nuances: "No other human activity represents this level of danger." The engineer, who previously worked at OpenAI, maintains that both labs have relaxed their safety protocols under pressure to keep advancing, and he places the catastrophic risk—that AI ends up wiping out humanity—at the end of this decade.

Seen that way, it would be the umpteenth apocalyptic prophecy from a sector addicted to hype. What changes things is who has come out to back him from inside.

The internal support that breaks the reassuring narrative

An Anthropic researcher, Evan Hubinger, responded to the case with a bluntness unusual in this business: he essentially says Coxon is right, that at the lab they truly believe AI could kill all humans, and that he puts that possibility above 10% in the next decade. He adds the most uncomfortable part: there is no plan to solve the alignment problem in a superintelligence, and there are no clear signs that they are on the right track.

In other words: this is not denounced by a talking head or an account seeking notoriety, but by someone who works on the problem from the inside and signs it with his full name. A probability greater than 10% of global extinction is not a figure to be dismissed with a shrug.

The Amodei precedent: those who warn and then found

It is worth recalling a detail that already peine. Anthropic's founder, Dario Amodei, previously worked at OpenAI and also warned about the dangers of AI. Then he set up his own AI lab. That parallel fuels the most repeated suspicion: existential risk discourse as a calling card to raise capital. If the pattern repeats, this episode would not be a warning, it would be a résumé.

It must be said with due caution: these are conjectures. No one has announced that Coxon is going to found anything, nor is there a known decision in that regard.

Marketing, bubble or technocracy?

Skeptical readings go in three directions and none is kind. The first is timing: it is noted that no one leaves Anthropic just before its IPO, so either he was invited to leave or he is paid better elsewhere. The killer AI narrative would function, from this perspective, as the new "peak oil": a narrative to keep living off the tale.

The second is money. It is argued that alarmism is the perfect alibi for the State to eventually step in and rescue the sector when the bubble bursts—the AI, real estate, and debt bubbles—with the explosion set for autumn. Just like in 2008: first the scare, then the bailout.

The third is deeper. If models do human work and elites no longer need us, the danger is not the machine, but who plugs it in. The analogy circulating is that of the horse and the internal combustion engine: the animal did not disappear because of the engine's malice. And the scenario being drawn—no privacy, no property, managed by a technocratic elite—can be summed up in one phrase: we will own nothing and be happy.

Who presses the button, the model or the owner?

Here is the real technical discussion. A missile launch system already requires protocols and multiple human approvals precisely because it is delicate. If an AI is connected, does anything change? The strongest argument holds that it does not: anything an AI does is first enabled by its programming, and poorly designed linear software can fail exactly the same way. The danger is not born from the machine's free will, it is born from where you place it and who takes the blame when it fails.

The anecdote that best illustrates the contrast is old: during the Cold War, a Soviet duty officer refused to return what his system presented as an imminent nuclear attack. It was a false alarm. He did not press the button because he chose to doubt. A system trained to maximize efficiency does not doubt.



The unsettling fact remains. The man who warns is 27 years old, the horizon he points to is the end of this decade, and the colleague who backs him puts the probability above 10%. It may all be smoke or the prelude to a funding round. It may also be the most honest thing said in the sector in a long time. And while it is being decided, the clock keeps ticking.

Summary of a discussion on Burbuja.info - Foro de economía, actualidad y política., translated from Spanish and reviewed before publication. Read the full discussion (87 replies).

More summaries

All summaries in English →

Back