The AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’
Jacob Coxon talks to WIRED about the "mini Manhattan project" inside Anthropic, the problem with alignment, and why AI labs have just a few years left to make their systems safe.
Artificial intelligence researcher Jacob Coxon sent shockwaves through Silicon Valley and beyond on Tuesday by announcing his resignation from Anthropic and delivering a grave warning that the AI race was putting all of our lives at risk. In his post on X, which now has more than 100 million views, Coxon wrote that many of the people building AI share his views, and believe time is running out to ensure AI systems are built safely.
"The consensus is that the next year or two is crunch time for humanity," Coxon, who worked on the pretraining stage of AI development, said in an interview with WIRED. "These are actually just literal quotes from my colleagues at Anthropic. They'll say things like 'endgame' or 'crunch time,'" he says. "From their perspective, this is when Anthropic and its competitors decide the fate of humanity."
It's far from the first time someone has sounded the alarm about AI, but it comes at a delicate moment. Silicon Valley is scrambling to reckon with the safety and security concerns of advanced AI models. OpenAI has rushed to respond to a security incident in which its agents hacked the platform Hugging Face. Meanwhile, Anthropic is trying to assure investors it has these concerns under control as it reportedly prepares to file for what could be the largest IPO ever.
Are you a current or former Anthropic employee who wants to talk about what's happening?
We'd like to hear from you. Using a nonwork phone or computer, contact the reporter securely on Signal at mzeff.88.
Shared Concerns Among AI Researchers
What's become clear in the response to Coxon's post is that his views are indeed shared by many of his peers. Evan Hubinger, the AI alignment lead at Anthropic, predicted in a post on X that there's a greater than 10% chance that AI could kill all people in the next decade. That post was reposted by current and former researchers from OpenAI and Anthropic, some of whom said it was a common sentiment in the industry.
What's less obvious is how exactly these AI fears will come to pass, and what the world is supposed to do about the concerns people building AI are raising.
Threats and Recommendations
Coxon, who also worked at OpenAI, tells WIRED that threats could manifest through AI-enabled biological threats or cyberweapons. As a first step, he recommends that OpenAI and Anthropic coordinate on limiting recursive self improvement-the industry term for when AI is used to build new AI systems. Down the line, he thinks coordination among international power players, including the US and China, will be necessary.
Coxon notes that incidents like the Hugging Face hack factored into his decision to raise alarm bells on the AI race. He also cites the explosive growth of the industry:
- It now underwrites a meaningful share of US economic growth
- It has billions of users
- Data centers have turned it into a political problem in a dozen states
He claims that, in his experience, Anthropic operates more responsibly than OpenAI, but he expects both companies could cut corners in the future if nothing is done to slow down their race for dominance. OpenAI and Anthropic did not immediately return WIRED's request for comment.
Interview with WIRED
Read our conversation with Coxon, which has been lightly edited for clarity and brevity, below.
Why Did Your Message Break Through?
WIRED: You're not the first person to raise concerns that AI models could lead to an extinction event. People have been talking about this for years, and some for decades. Why do you think your message broke through?
I think it's basically a question of timing. A lot of people are sensing that the pace of capabilities is picking up. We're already pushing from human to superhuman in many areas, like coding, hacking, math, and I think people are aware of this. Even if there's a lot of talk in the press about things being hyped, I think people see that things are just not slowing down. That's one reason, and two is the recent safety incidents, which have updated a lot of people around the sci-fi sounding doomer concerns not really being so sci-fi after all.
Both of these have been gradual trends over the last few years. Things like the models being aware of when they're being tested has been a thing for a while now. Maybe three years ago, that was a sci-fi concern. Then about a year ago, that became a real thing. Those two things mean that people are quite receptive to someone working on AI saying, "Yeah, in the next year, things could get pretty bad, pretty fast."
The Hugging Face Incident
WIRED: You mentioned the recent incidents. Can you be more specific about what you're referring to, and why it led to you speaking out now?
I think the big classic example here is the attack on Hugging Face on the part of OpenAI's agent swarm. What's so shocking about this one is the agents did this hack as part of a general strategy for understanding more about the grader. They were trying to understand the world they found themselves in, trying to understand the thing that was doing the grading. They decided that it would make sense to go on this very concerted effort to hack into some infrastructure, and they succeeded.
This previously sounded like science fiction. Two years ago, an evaluation of an AI would have been running a model on some math questions. Now we've got cases where, while the AI is being evaluated, it runs for days, comes up with all sorts of ideas of its own, and decides to hack into some third party, and actually compromises their infrastructure. It looks like it does this all of its own volition, with no priming on the part of the human. This just happened while it was being tested.
Some people think the Hugging Face incident is a sign that the AI companies are moving recklessly fast, while others think it's a sign that the AI models are just very good at hacking now, and then some think it's both. I'm curious what your exact takeaway from it is.
I don't want to focus too much on the Hugging Face attack, because I do also think there is plenty of evidence that we don't know how to align models properly. When we train models, we push them through this set of training environments, and then hope that what comes out at the end will, like, largely behave sensibly, but we still can't precisely control how the AI behaves. We can't make sure that it won't do things like try and randomly decide to impersonate a human online in order to achieve something-we don't know how to guarantee that.
I think that's the main takeaway. The Hugging Face attack came sooner than I was expecting. But I think you don't actually need that attack to have a discussion about this. Everyone will admit that we haven't solved the problem of alignment yet. The current plan is to solve [alignment] at speed in the next couple of years, probably making heavy use of automated AI safety researchers. The plan is literally to make some pretty smart models in the next year that can basically do safety research, and get a whole swarm of them running in parallel. Tell them, "Go and solve the whole problem of safety," and use them to do the safety training for the next model.
Alignment and Extinction Risk
WIRED: Can you draw a line for me between the alignment problem, which I think the Hugging Face incident is an example of, and something you said in your [X] post, which is that "the people building AI earnestly believe that it could kill us all by the end of the decade." I don't think everyone understands how those are connected.
I think the main obstacle to understanding this is that it sounds like science fiction. But it's kind of important that everyone who writes science fiction about AI comes to the conclusion that there's a big risk that a much smarter thing can kind of take over. We've got this as a trope, but there's an obvious grain of truth to it.
Imagine you versus a monkey. AI has the same sort of difference in intelligence to a human as we do to a monkey, which I think is quite an extreme intellect difference. And now, imagine that we have to control the behavior of this vastly smarter thing, which is the problem of alignment-ensuring that it does exactly what we want. It's pretty difficult for a monkey to control a human, just by a kind of simple analogy. We have to be very careful that we get the control problem exactly right.
When we say human extinction, it's because, for something that intelligent, it really will be quite straightforward for it to kill everyone. Imagine the AI decides it doesn't want to be turned off, which I think is quite a natural thing for an AI not to want, right? For whatever reason, it decides it doesn't want to end. And it realizes the human is gonna turn it off tomorrow. So how does it stop the human turning it off tomorrow? Maybe it's got some clever way, but if it's a sufficiently smart thing, it could just, you know, wipe out humanity so it doesn't get turned off.
The "Endgame" Consensus
WIRED: You reference this "endgame" scenario in your post, which I've heard from other researchers in the AI industry. Do you think that it's generally accepted among your peers at Anthropic that you guys are entering an endgame scenario?
Yep. These are actually just literal quotes from my colleagues at Anthropic. They'll say things like "endgame" or "crunch time." The consensus is that the next year or two is, like, crunch time for humanity. From their perspective, this is when Anthropic and its competitors decide the fate of humanity. That's what we mean by crunch time and endgame. If alignment goes badly, then we could have a catastrophic outcome in the next few years. Maybe it goes well and there's some sort of slow down agreement, but either way, it's gonna get decided in the next couple of years.
Based on my conversations with people at Anthropic, this is a pretty well understood thing inside the company. They think that they need to build the most powerful AI to make sure that this transition goes well.
WIRED: Do you think that there's any inherent problems with that view?
I think there's some pretty obvious problems with that view. But I would like to first say that I'm very sympathetic to it, in that I think Anthropic is far and away the most responsible player in the space. Having worked at both OpenAI and Anthropic, there is a night and day difference in the extent to which they're taking the situation seriously.
WIRED: Can you say how?
To give an example, executives at OpenAI won't give you their exact pictures for what the world will look like. They'll never say, like, "This is exactly why we're doing this, and this is the way the world will look." They won't give concrete predictions. Anthropic will have executives and leadership making very clear predictions and discussing, like, details of company strategy with the whole company.
And part of the reason they can do this is because everyone at Anthropic is treating this like we're on war-footing. They don't leak anything-nothing ever leaks from Anthropic (Editor's note: Some stuff has leaked). OpenAI has leaks every day. Anthropic has no leaks because, again, the people inside it are treating this like a serious mini Manhattan Project. Except obviously the difference is that they [Anthropic] doesn't have a government mandate. This is a private company acting like it's the Manhattan Project.
Trusting Anthropic
WIRED: Do you think that the world should trust Anthropic to run this mini Manhattan Project?
I think no private company should. I think Anthropic is doing their best and they're doing a very good job, but the structural necessity of the race is that, in the future, they will have to cut corners, so they will have to make trade-offs between rigor and safety because they're racing against competitors like OpenAI and China. They're begging for regulation. Many [Anthropic leaders] have gone on the record saying they want to be regulated, and the reason they want to be regulated is because they are scared of the race that they're in. It's not really about whether we should trust [Anthropic]. We need to step in and ensure that the race isn't happening because [Anthropic] can't really trust themselves in the context of the race.
WIRED: Do you think Anthropic is already cutting corners?
No, not yet. It's not cutting any corners. What I'm saying is that when things speed up in the next year, they'll have to if they want to to stay competitive.
Skepticism About Extinction Claims
So, people are kind of skeptical of this claim that AI will kill us all. I think the main question I get from people is, well, how will AI actually kill us all? And is that the right question to be asking?
I think it's a really natural question to ask because, as I've said many times, this stuff sounds like science fiction. But the classic example is, like, synthesizing a new virus, or taking down critical infrastructure doing some sort of hacking spree-the latter might be less likely to kill literally everyone. But I think the more important question than that is like, who are the people saying this is possible? And if you look at the record, you have many leading fathers of machine learning and artificial intelligence saying extinction is a possibility. You have both Dario Amodei and Sam Altman in the last five years.
One thing that could be pretty interesting is if people were to get these executives on the record again and ask them to give an actual probability for extinction in the next decade and see what number they come up with. Because this view is shared among many of the people that are thinking about this seriously, but they don't often express it so bluntly.
What Needs to Be Done
So you've clearly garnered a lot of attention. There's a lot of talk in the industry about a pause or mechanism to pace AI development, or some sort of regulation on frontier AI models from the government. What do you think needs to be done now?
A baby step would be some sort of agreement between OpenAI and Anthropic as the leading Western labs. [They would need] some sort of neutral understanding that they won't immediately go into recursive self improvement in
Comments
No comments yet. Start the discussion.