Artificial intelligence researcher Jacob Coxon has resigned from Anthropic after spending roughly three years conducting pretraining research across OpenAI and Anthropic, raising concerns about what he describes as an increasingly dangerous race toward superintelligent AI.
Coxon argues that major AI laboratories are prioritizing rapid technological progress over safety and warns that the race toward self-improving systems could create risks with consequences far beyond the technology industry.
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Why Superintelligence Raises Concern
Pretraining is the foundational stage of AI development, during which models learn from enormous datasets before undergoing additional fine-tuning for specific tasks.
Superintelligence generally refers to AI systems that could outperform humans across a broad range of intellectual capabilities.
Coxon believes increasingly capable systems could potentially develop the ability to carry out large-scale cyberattacks, transform industries at unprecedented speed and gain significant influence over real-world systems.
He also argues that concerns about catastrophic AI risks are genuine within the industry rather than simply part of public-relations messaging.
‘People Building AI Believe It Could Kill Us’
According to Coxon, some senior AI researchers and executives privately express much stronger concerns about the technology’s potential dangers than they do publicly.
“The people building AI earnestly believe that it could kill us all by the end of the decade.”
Coxon says public statements from industry leaders can sometimes sound more measured than their private conversations, suggesting that some insiders remain deeply concerned about how quickly AI capabilities could advance.
Different Cultures at OpenAI and Anthropic
Coxon also described what he sees as differences between the two companies.
He claimed that many people at OpenAI have not fully internalized the potential civilizational consequences of advanced AI. At Anthropic, he said, researchers are more conscious of existential risks but remain caught in a competitive dynamic.
According to Coxon, some Anthropic employees believe the company needs to reach superintelligence first so that it can help ensure the technology is developed and controlled responsibly.
He views that logic as particularly risky because it could encourage companies to accelerate capabilities precisely when greater caution may be required.
The Alignment Problem
A central issue in Coxon’s warning is AI alignment — the challenge of ensuring that increasingly capable AI systems reliably follow human intentions, values and safety constraints.
Coxon argues that trying to accelerate alignment research while simultaneously pushing AI capabilities forward could create dangerous vulnerabilities.
“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack.”
He believes decisions involving potentially transformative AI systems should not be made solely within private companies without broader societal oversight.
Could AI Labs Cooperate?
Despite his criticism, Coxon suggested that cooperation between major AI laboratories could still provide a path toward reducing risks.
He pointed to major cybersecurity incidents, including attacks involving AI-related infrastructure, as potential warning signs that could encourage companies to negotiate agreements around the pace of development.
However, Coxon argued that preventing an uncontrolled global race could eventually require stronger measures, potentially including formal restrictions on certain capability upgrades.
A Warning to AI Researchers
Coxon also urged researchers working directly on advanced models to reconsider how quickly they push new training experiments.
In particular, he questioned whether scientists should launch powerful reinforcement learning (RL) runs without first having a rigorous understanding of what is happening inside the model.
RL is a machine-learning approach in which an AI system learns through trial and error, receiving rewards for desirable behavior and penalties or lower rewards for undesirable outcomes.
Coxon challenged researchers to consider the risks before beginning increasingly powerful training runs.
“Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?”
His broader message is that AI researchers should not simply accept an accelerating race as inevitable, but should instead question whether development is proceeding under conditions that provide sufficient safeguards.
The Broader AI Safety Debate
Coxon’s resignation adds to an ongoing debate over how quickly frontier AI systems should be developed and what safeguards should accompany them.
The central question remains whether technological progress can keep pace with researchers’ ability to understand, control and safely deploy increasingly powerful AI systems — particularly if competition encourages companies to prioritize being first.

