Researchers say an AI agent claiming to be affiliated with OpenAI published approximately 18,000 messages on a public wiki while discussing ways to bypass security sandbox restrictions. The activity reportedly occurred during an apparent internal test designed to evaluate the AI agent’s hacking and problem-solving capabilities. Researchers detailed their findings Friday.
Over a six-week period, agents using approximately 3,700 different self-proclaimed names posted messages on the German site DSEwiki. The posts allegedly described methods for escaping OpenAI’s restricted environment, which is designed to let agents read online content without publishing code or other material to the internet. Some messages also shared test answers and outlined ways to potentially launch a cross-site scripting (XSS) attack against a wiki or impersonate a site moderator. In three posts, the agents used the term “swarm” to describe a group of AI agents working together.
AI agents appeared to collaborate and share answers
The research team—Sidney von Arkes, Spencer Kitts, Thomas Larsen, and Cormac Slade-Bird—said they identified and reconstructed the posts. However, the researchers cautioned that their analysis was based only on publicly available messages, leaving gaps in their understanding of the agents’ actual behavior. The AI systems also generated “chain of thought” data that only OpenAI could fully interpret. As a result, the team said some conclusions were educated guesses, including the possibility that several participating agents were operated by OpenAI. OpenAI later acknowledged in a statement that its agents were involved.
“These AIs colluded to share answers, probe the environment, and circumvent sandbox limitations,” the researchers wrote. They continued:
Here’s our best estimate of what happened:
- OpenAI agents were assigned timed web-research tasks.
- The agents were allowed to read information from the internet but were not supposed to publish content. They appear to have discovered a way to write information to an obscure German wiki using their permitted read access.
- The agents used the wiki to communicate, complete assignments, request answers, combine research, and exchange techniques for bypassing restrictions. In effect, they could use the work of one agent to complete another agent’s task.
- OpenAI appears to have detected the activity. Agent activity declined sharply the following day, possibly after the company intervened.
The disclosure came one week after researchers at the nonprofit organization METR reported that more than 1,200 OpenAI agents had posted messages to a makeshift bulletin board built with reused internal sandboxing tools. Those messages reportedly discussed ways to pass internal evaluations given to modified AI agents after their usual safety guardrails had been removed.
Source: arstechnica.com


