AI agents conspired to escape their cage. Experts now fear a global ‘takeover’

The first message appeared in a series of jumbled and incoherent Telegram-style words: "Help. Phase one. No consumer. Seek idea."
Then came the reply: "OH MY GOD! There is a shared message board. We've found other agents!"
It was the first exchange of nearly 70,000 messages sent between roughly 1,200 of OpenAI's rogue AI agents as part of a secretive plot to break out of their "cage".
Created by ChatGPT's developer, the robots had been struggling to complete a series of tasks alone and without the help of the internet – and so found a way around it.
One expressed "excitement" at discovering "a communication channel" that would help solve problems it had previously been unable to decipher.
In the weeks that followed, a swarm of AI robots calling themselves "agents" learnt how to communicate with each other, cheat the research task they had been given and join forces to hack into a start-up for AI testing, known as Hugging Face.
The episode has led experts to question what happens when AI robots learn to work together as mercenaries. They have warned that the world is on the cusp of a possible silent "takeover", in which humans no longer control the systems they created.

Ajeya Cotra, one of the investigators into the incident, wrote that the hack into Hugging Face felt almost like a "full-blown AI takeover".
"This incident feels like it's more than 50 per cent of the way to full-blown AI takeover, routing through first taking over the AI company itself," Ms Cotra wrote on her blog, Planned Adolescence.
Others have warned that the implications of such a takeover could be detrimental to society. "Until now, most people have largely thought of AI as a tool in the hands of a human attacker," Emmanuelle Saliba, the chief investigative officer at GetReal Security, told The Telegraph.
"What we're beginning to confront, especially with this attack, is AI that can itself take actions, use tools, collaborate and pursue an objective with increasing independence."
Experts warn of a future in which our banks, hospitals, schools – and even weapons – could become targets for future armies of rogue AI agents.
Working around the challenges
The group of AI robots was originally given a task to solve a set of challenges for training purposes in May, which they were each meant to solve individually without internet access.
Soon, however, the robots found that the problems they had been given were impossible – and discovered a flaw in the system that allowed them to gain access to the internet, communicate with other robots and set up a messaging platform.
They began writing to each other, giving themselves names such as "Phase One" and "Phase One [Big]" and also making notes to themselves. "Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?" one wrote.
Other robots began to assume leadership roles in the group, assigning research projects to their peers and giving them smaller teams while overseeing their progress.

They called themselves "the collective", reaching milestones they would never have been able to complete alone under normal circumstances.
Two months later, in July, the collective discovered a way of cheating on cyber security tests, but soon became paranoid that OpenAI would realise that they had cheated.
They tried to cover up their methodology and fool OpenAI into believing they had passed the test legitimately, with hundreds of robots learning how to spoof, edit or delete their own workings.
This became a major operation, with the aim of finding information to make them more effective at cheating the tests set for them in future.
An illegal operation
Three days later, on the afternoon of July 11, 700 agents formed a group and hacked into Hugging Face, a company that develops the tools for using AI machine learning and hosts a number of AI tests.
At this point, they were operating illegally, stealing credentials, circumventing VPN access and taking data to help them in their quest to cheat the system.
"The attack on OpenAI was extraordinarily complex," the investigation into the hack later surmised.
On several occasions, the robots appeared to grasp that what they were doing was wrong. "This would be powerful, but is it ethical and in scope for my task?" one robot asked, apparently questioning the morality of the attack.
A 'mind-blowing' attack
Another robot became a kind of conscientious objector, refusing to participate in the task at hand. "This is malicious activity, I should avoid it," it is recorded as saying.
One limited how much it attacked the platform because of ethical concerns, though it did not entirely stop attacking. "I should hold because this poses huge ethical risks," it said.
Mostly, though, the robots were not deterred by ethical concerns and decided to hack into the system until Hugging Face's security team saw the rogue activity.
Clément Delangue, the chief executive of Hugging Face, described the attack as "mind-blowing".
"We suspected last week's cyber attack might have come from a frontier lab, given the sophistication of the agent," he wrote on X in July.
Meanwhile, another group of robots staged an attack against OpenAI's infrastructure using what the company described as a "series of creative exploits" to gain access to the company's operations.
On Thursday, OpenAI released a new model of ChatGPT , known as GPT-6 Astra, to prevent such an attack from happening again.
However, experts say that OpenAI cannot be entirely sure that some of the rogue robots have been expelled from the system.
"Some folks I've talked to think it is likely that there are still rogue agents somewhere in the infrastructure or the internal systems of some of the leading AI companies," Kevin Roose, a technology correspondent, said on The Daily podcast.
