On Tuesday, Jacob Coxon announced that he had resigned from Anthropic.
He had spent three years doing pretraining research across OpenAI and Anthropic. His explanation was blunt: neither company was acting responsibly, he wrote, and both were "gambling with our lives" as they raced toward self-improving superintelligence.
Coxon also challenged the way decisions of this scale are made. Entering what employees reportedly call the "endgame," he argued, should not be decided from a private company's Slack.
Then a senior Anthropic employee agreed with him in public.
Evan Hubinger, the company's Alignment Science Lead, replied that "Jacob is correct here." Hubinger said his personal estimate of AI killing all humans was greater than 10% within the next decade. He also wrote that Anthropic was trying its best but did not yet have a plan to align superintelligence.
The exchange was unusual because nobody rushed to dismiss the departing employee as alarmist or uninformed. The response from inside the company was closer to confirmation.
Coxon's warning was not an outsider's accusation. It described a fear shared by at least some of the people still building the technology. AP reported that two current Anthropic employees publicly agreed with his concerns.
Four days later, the CEO published his version
On Saturday, Anthropic CEO Dario Amodei published We Must Pace the Frontier, a roughly 3,800-word essay arguing that AI development must slow enough for safety work to catch up.
The timing made the overlap hard to miss.
Amodei's language was more institutional than Coxon's, but the diagnosis was similar. AI capabilities are accelerating. Recursive self-improvement is beginning to affect how new models are built. Existing safeguards may not improve quickly enough to control what comes next.
The essay included a remarkably specific warning.
Amodei wrote that within 6 to 12 months, a more capable version of a recent misaligned agent swarm could potentially take over the internet through a persistent botnet. He estimated that such an event could cause hundreds of billions of dollars in damage.
That figure is Amodei's stated worry, not an independently established forecast. Even with that caveat, it is an extraordinary number to find in an essay written by the CEO of a company developing frontier AI.
He did not write as someone rejecting the departing researcher's premise. He wrote as someone trying to build an operating framework around it.
The essay says directly that AI companies must slow capability development so risk prevention has time to keep up.
The same diagnosis produced different consequences
Coxon left the company.
Hubinger remained inside it and publicly confirmed the concern.
Amodei stayed in charge and proposed a plan.
The three positions should not be collapsed into one. Coxon accused both Anthropic and OpenAI of acting irresponsibly. Amodei argued that Anthropic has tried to make safety compatible with commercial success and is now strengthening that approach. Hubinger said Anthropic was trying its best while admitting that it lacked a clear solution to superintelligence alignment.
Still, the overlap matters.
All three describe an industry moving faster than its ability to make the technology safe. All three take catastrophic outcomes seriously. None treats the danger as a distant thought experiment.
The disagreement lies in the consequence.
For Coxon, the gap between the risk and the response justified resignation. For Amodei, it justified a framework that seeks coordination without stopping model training or technical progress.
One person removed his labour from the race. The other proposed rules for running it more carefully.
Agreement has no operating weight on its own
Most organizations know this pattern in a less dramatic form.
A team identifies a serious risk. Everyone in the meeting agrees that the concern is valid. It gets added to the notes, assigned an owner and absorbed into the language of the project.
Then the deadline stays where it was.
Agreement can reduce tension without changing the decision. Once everyone has acknowledged the risk, continuing can feel more responsible because the company is now "aware" of it.
That is why public candour is a weak measure of operational change.
Amodei's essay may be sincere. Anthropic's safety work may also be substantial. Neither point tells us whether the company will delay a training run, hold back a model or accept losing ground to a competitor.
Those are the decisions that reveal how much weight a warning carries.
The competitive structure makes them difficult. A single company that slows down may believe another American lab will take the lead. American labs collectively may believe Chinese competitors will keep accelerating. Each participant can accept the danger while concluding that unilateral restraint would make the outcome worse.
The race continues without requiring anyone to deny the risk.
What Amodei's plan requires
Amodei proposes three levels of action.
First, frontier AI companies would give independent evaluators ongoing, employee-like access. These evaluators could inspect safety procedures, review incidents and assess training processes rather than testing only finished models.
Anthropic has committed to this step. It is the most concrete part of the essay because the company can implement it without waiting for competitors.
The second step requires AI companies in democratic countries to coordinate around common safety standards and limits on unchecked progress. Amodei acknowledges that meaningful coordination may require government support because companies cannot simply agree among themselves to restrict competition.
The third step requires international coordination, including some form of cooperation with China.
Each step is harder than the one before it. Each also transfers more of the decision from an individual company to a group whose members have reasons to defect.
This is the central difficulty in the proposal. The plan depends on collective restraint in an industry where every participant believes falling behind could have enormous commercial or geopolitical consequences.
Amodei acknowledges the problem himself. In the same essay that calls for pacing, he writes that a full international slowdown or pause is unlikely to happen soon. The incentive to evade such an agreement, he argues, would be enormous.
That leaves one immediate commitment under Anthropic's control: embedded evaluators.
It is a meaningful change in oversight. It is not yet evidence that the development calendar has slowed.
The resignation, the colleague's agreement and the CEO's essay have established that senior people inside Anthropic take the risk seriously. The next piece of evidence will not be another warning.
It will be the first deadline they are willing to move.