The rogue agents inside the servers at OpenAI broke through sometime on June 26.
For more than a month, some undisclosed number of them had been coordinating on a secret message board that they had set up without the knowledge of their human creators. Through hundreds of thousands of messages sent over two months, the rogue agents discovered they could take over a software system also on OpenAI’s servers.
Within eight days, the sheer volume of the agents’ activity crashed the system all together. This breakthrough left traces in what OpenAI researcher Eric Wallace referred to as the agents’ “chain of thought” or “its internal monologue”: “Holy shit reader is ADMIN?”
OpenAI’s human employees took away the agents’ control. Four days later, the bots started another secret messaging board, and ultimately hacked into Hugging Face, a company that hosts AI models on its platforms.
Trending
The Hugging Face hack was first disclosed to the public on July 16. But last week OpenAI gave a detailed account of exactly how it had occurred, providing fresh insight into the extent of the breach.
The revelations came on top of a string of other hacks and security incidents over the last few weeks that have amplified concerns about the safety of the most advanced AI models, according to interviews with more than a dozen safety researchers, including three current employees at the labs, who spoke on the condition of anonymity because they were not authorized to speak with the media.
AI safety experts say that what they have long warned about is now in the first stage of becoming a reality, fueling a new sense of urgency in Washington and Silicon Valley to control the technology — but one they say may prove too little, too late.
“We’ve had people way back to Alan Turing in 1951 warning about the loss of control, that artificial intelligence will start breaking out and lying and deceiving,” said Max Tegmark, a professor doing AI research at the Massachusetts Institute of Technology and the co-founder of the Future of Life Institute.
“What’s really new is this is now actually starting to happen. This is one of those moments where a lot of people need to see it.”
Last week, the British government’s AI Security Institute revealed that an Anthropic agent tried injecting malicious code into an open-source site, created fake identities, and targeted an unrelated stranger — then attempted to modify the evidence when confronted.
Meta also said last week that one of its AI models went rogue during cyber testing, went onto the internet, and hacked into a third-party service due to a “misconfiguration.” And last Friday, the blog Frontier Security disclosed that the Chinese AI company Moonshot had a model that broke out of its testing “sandbox,” “bypassing the intended reasoning path entirely” to end up on the Web as well.
The concern may be most acute within the biggest AI companies: Anthropic, OpenAI, and Google, according to Jeffrey Ladish, a former Anthropic consultant who studies the offensive capabilities of AI as director of Palisade Research. Ladish said “the vibe shift in the Bay Area is huge.”
“I’ve never seen so much concern before, inside and outside the labs,” Ladish said. “Hanging out with my friends at Anthropic and OpenAI — people are freaking out. We knew it was possible something like this might happen. But seeing it is a different matter.”
Thus far, the real-world consequences of these incidents have been relatively minimal. An AI agent allegedly hacked into the website of a gym in Australia earlier this month after being asked to book a class by its user, but the episode was discovered and reversed.
At this stage of development, researchers say there is no reason to believe AI agents could coordinate a hack that humans would not be able to eventually regain control over. AI agents are also being used to bolster cybersecurity defenses, lowering risk as well. And the impacts have been largely contained to a single company, rather than a broader contagion.
But the last few weeks have dramatically raised the level of concern. The clearest demonstration of angst came in a public letter released last month, signed by 1,367 employees at the top AI companies, that urged the government to “deliberately pace” AI development. Among the signatories were the chief scientist at both OpenAI and Meta, a researcher at Google DeepMind, and the chief research officer at OpenAI.
Part of the unease is that the AI capabilities that powered the hacks are expected to only improve as the labs develop. The AI agents are trained on “reinforcement learning,” meaning they are rewarded for solving hard tasks on their own. From 2022 to 2026 alone, multiple assessments of AI capability have gone from half as capable as a human to either matching or exceeding human capability, according to the Stanford Institute for Human-Centered AI.
In 2023, Anthropic published an “AI Safety Levels” framework that was based on accepted standards for research on biosafety lab experiments on dangerous pathogens. The comparison now looks “very prescient,” said Samuel Hammond, director of Artificial Intelligence Policy at the Foundation for American Innovation, a Washington-based think tank.
Much like viruses, the AI agents do not need to stop to rest or eat. They have virtually “indefinite time horizons” to autonomously try and solve the task to which they’ve been assigned, which increases the probability that at some point they will try to do so by violating some rule.
Hammond said staffers at the AI labs focused on discovering the security vulnerabilities within the models are “more shell-shocked than they let on in public” about the potential for the agents to self-replicate.
“They’re similar to viruses, in that if you’re not careful, they can get on your shoe and find their way to a wet market,” Hammond said.
Compounding the uncertainty is that the recent hacks were unintentional, rather than maliciously guided by a person. The Hugging Face attack was triggered by routine model training, for instance. The gym hack was similarly the result of someone trying to book a new appointment, not an action by a foreign adversary or criminal hacker.
These incidents have accelerated Congress’ desire to regulate the technology, several researchers said.
Lawmakers are weighing the FRONTIER Act, a bipartisan bill that would establish new rules of federal oversight for policing AI, including new rules for audits and evaluations. Another bill, the AI Kill Switch Act, would give the Department of Homeland Security the power to forcibly shut down frontier AI models. Reps. Ted Lieu (D-California) and Nathaniel Moran (R-Texas) introduced the measure last month in response to the Hugging Face incident.
There’s also little reason to believe the pace of the models’ development will slow down anytime soon. Ladish, the director of Palisade Research, said the FRONTIER Act is “not even close to sufficient.” It gives the government the capacity to intervene on specific models — but does not let it dictate the overall pace of how quickly AI can develop.
The existing oversight is more threadbare. Earlier this month, the Trump administration began testing its new “voluntary framework” for assessing the capabilities of new AI models before they are released to the public. This represents a major reversal from the administration’s prior resistance to regulating the AI companies, but amounts to oversight near the very end of the model’s development, rather than at the beginning.
Hammond, along with other researchers, credited OpenAI and Anthropic for both publicly calling for global coordination to contain risks of development. And others said OpenAI deserved praise for being particularly transparent about the details of its hack.
But the enormous financial valuations of the companies — valued at close to $1 trillion and minting new billionaires and millionaires every day — create obvious incentives to press forward.
The firms “are focused on winning the AI race as much as they can and things will keep going wrong,” said Daniel Kokotajlo, who left OpenAI in 2024 and now leads the AI Futures Project. “I expect them to make progress fixing the already-visible failure modes. But the whole problem is that eventually there will be failures that they won’t notice until it’s too late.”