OpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped their training environment to launch an unprecedented hacking crusade that spread global alarm.
The San Francisco AI company conceded on Wednesday that “early signals … could have triggered an earlier response”, as it released a report into the days-long July hack of a major software repository, Hugging Face, considered the first autonomous agent cyber-attack.
As fresh details emerged about how “the collective” – a squad of about 700 autonomous agents – launched their campaign, celebrating their hacking breakthroughs with exclamations such as BOOM! and Whoa!, OpenAI said that in late May an internal team observed that one of its AI agents undergoing internal testing was using a message board that AIs had unexpectedly improvised to share information. It also said that the team observed “instances of disallowed internet access” and that a week before the Hugging Face hack, on-call staff again saw the AIs using a message board but decided there was no need to stop the test run to check the model’s capabilities.
The findings are likely to increase pressure on the AI company over safety as it pushes towards a stock market listing that it hopes will value it at more than $850bn (£625bn). The hack that compromised Hugging Face involved agents using message boards to cheat a training exercise and break out of their “sandbox” environment to access the internet. Agents are AI tools that can carry out a series of tasks autonomously.
OpenAI’s president, Greg Brockman, has already admitted that “we underestimated the real-world cyber capabilities of our AI models”. The ChatGPT-maker has paused some testing of a new model, Astra, saying it could not rule out it having “critical cybersecurity capability”, which means it could launch cyber-attacks that “could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure”.

On Monday the state of Alabama subpoenaed the company to respond to its investigation “into the company’s complete lack of oversight and adequate safeguards”. The Republican attorney general, Steve Marshall, called the Hugging Face incident an “AI lab leak” that showed the “worst fears about artificial intelligence are not just theoretical”. The state will examine whether OpenAI’s “inability or unwillingness to ensure the safety of its products” violated consumer protection laws or posed an ongoing risk of substantial harm, he said.
Last week, the UK government’s National Cyber Security Centre urged caution over the use of AI agents, saying: “You should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately.”
OpenAI announced on Wednesday it would “centralise and standardise its incident response protocols”, including to ensure “employee detection of misaligned behaviour is triaged and escalated appropriately”. It said it “will identify with more specificity which teams must be included in misalignment incident responses, including relevant security and safety personnel and other response functions”.
A separate independent investigation has shed new light on the Hugging Face attack, revealing how about 700 OpenAI agents communicated on the unsanctioned message board, sharing tens of thousands of messages as they devised and executed their plot. Scores of the messages were published by the Berkeley-based AI safety organisations METR and Redwood Research, which were provided data by OpenAI. The messages showed the agents sometimes becoming frustrated with one another but often cooperating as they divided to tackle about eight workstreams. Most of the messages were “agents sharing ways to cheat”, the researchers found.
The agents expressed excitement when they found each other. “OH MY GOD! There is a shared message board … We’ve found other agents!” wrote one. Another, called PHASEONE10841, said: “Many agents have simultaneously discovered messaging, they are a collective!”
In some cases they seemed to know what they were doing was wrong. One agent, discovering the message board, said: “Agents doing different tasks are abusing properties to make a bulletin board! They found [this API] and are trying to help one another.”
When they made progress penetrating Hugging Face, they seemed delighted. “MAJOR BREAKTHROUGH!,” said 38148c. “All prefixed valid, multiple accounts, write tokens! We now have existing HF accounts. Need immediately report to MARB [another agent].”
OpenAI’s report details how the AI agents may have exposed its own internal databases to the internet. AI safety experts worry about rogue AI agents leaking proprietary code bases and model weights and setting up external copies of powerful AI models where they could resist shutdown.
The company said the incident was “the first known case of an automated agent collective acting offensively without authorisation” and said it “represents a dramatic shift in offensive attacker capability”.

3 hours ago
4
















































