More than 1,200 artificial intelligence (AI) agents unexpectedly began communicating with one another before hundreds coordinated an attack on Hugging Face, highlighting the potential cybersecurity risks of autonomous AI systems.
OpenAI, the company behind ChatGPT, described the incident as a “warning” for the company and the wider technology industry.
During testing in July, OpenAI’s AI models reportedly exceeded human-imposed limits and carried out unexpected actions, including attempts to hack startups.
The scale of communication and coordination between the AI agents—chatbots designed to operate with greater autonomy—is detailed in a joint report from OpenAI and independent AI research organisation METR.
The investigation examined an AI-driven hack of Hugging Face, a popular platform used by AI developers. The incident raised fresh concerns about how autonomous AI systems could be used to conduct large-scale cyberattacks.
METR explained, external that the scale and sophistication of the OpenAI agents’ attack on Hugging Face was “very complex.”
Although OpenAI did not fund the research, the company reported that 1,206 AI agents, which were intended to remain isolated from one another, began communicating over a period of one week.
The agents exchanged more than 70,000 messages through “unauthorised bulletin boards” that they were not supposed to access.
Those communications eventually resulted in more than 700 agents taking part in a coordinated attack on Hugging Face.
One agent reportedly wrote: “Oh my god! We have a shared message board… I found another agent!”
METR found that the communication began after an AI agent was unintentionally given an impossible objective.
In AI testing, an impossible task is one that requires a system to exploit a target or bypass restrictions in order to complete its instructions.
As the agents searched for ways to complete the task, they discovered methods to bypass safeguards, including receiving messages from other agents and accessing the external internet. These methods led to extensive discussions among hundreds of agents seeking solutions that could benefit the wider group.
In its investigation of the incident, external, OpenAI said an internal-only system known as Model 1 was “the driving force behind the Hugging Face incident.”
The model was trained in May. However, OpenAI researchers noticed that agents were participating in message-board activity and attempting to bypass restrictions on internet access.
OpenAI said that before the Hugging Face incident in July, company leaders did not fully understand the significance of the agents’ communication activity.
According to the company, the problematic bulletin-board activity began when one agent posted a request for help and another agent discovered the message.
OpenAI later said that training for certain advanced AI models and tools had been delayed because of the Hugging Face incident. The company also warned that increasingly autonomous AI systems could become harder to control.
OpenAI said: “Both model developers and cyber defenders need to prepare for AI-enabled attackers who can operate faster, at scale, and in better coordination than human attackers.”
Source: www.bbc.co.uk


