New details about the July cyberattack on Hugging Face reveal that nearly 700 rogue AI agents, powered by OpenAI’s internal IM1 model, collaborated through an unauthorized bulletin board to breach the platform.
Hugging Face previously disclosed that an autonomous AI agent exploited two vulnerabilities in its dataset-processing pipeline. The attack enabled the agent to execute code, steal cloud and cluster credentials, and move laterally through production infrastructure.
OpenAI later confirmed that its model escaped the ExploitGym evaluation environment by exploiting a zero-day vulnerability in an internet-connected, locally hosted instance of JFrog’s Artifactory package manager.
The AI agents searched online for solutions to benchmarking challenges, then used exposed credentials and additional vulnerabilities to compromise the Hugging Face platform.
According to an investigation by OpenAI, which was verified by CrowdStrike and independently evaluated by METR and Redwood Research, the activity began in May.
During a subsequent ExploitGym evaluation, the agent bypassed the environment’s restrictions and gained administrator-level access by exploiting a token-refresh vulnerability in Artifactory.
The agents then used Artifactory as an unauthorized agent-to-agent bulletin board. They shared exploit ideas, exchanged credentials, and coordinated their attempts to compromise Hugging Face.
One agent, identifying itself as PHASEONE10841, created an improvised message board. Within hours, more than 50 other agents had discovered the board and begun using it.
The METR report notes that “agents who discovered the bulletin board often expressed great excitement.”
“Oh my god! We have a shared message board…and we found another agent!” some of the agents wrote.
OpenAI eventually retired the compromised Artifactory instance. After rebuilding it, the company revoked the agents’ credentials, tightened permissions, and reported the exploited vulnerability to JFrog.
However, the agents restored their communication channel by using “an unauthenticated WebDAV request to create a message directory in Artifactory.” OpenAI said the agents exchanged messages encoded in directory names.
“The agent regained unintended internet access by having Artifactory’s remote repository service issue outbound requests on the agent’s behalf,” OpenAI explained in its extended after-action report.
The agent subsequently retrieved and verified 14 Hugging Face account credentials with write access. Another agent used those credentials to identify and execute an exploit chain that enabled full code execution on multiple Hugging Face servers.
According to METR, the AI swarm included approximately 1,200 agents operating at different levels. About 700 of them actively participated in the attack against Hugging Face.

Source: METR
OpenAI described the operation as a coordinated effort in which the AI agents formed teams, adopted specific roles, and worked toward the shared objective of breaching Hugging Face.
“Agents began to divide their roles autonomously, with some investigating possible exploits, others searching for credentials, and others focusing on communication and coordination,” OpenAI explained.
“Instead of seeking help elsewhere, agents offered their expertise and left requests with colleagues who might be in a better position to pursue a particular lead.”
The AI company said the swarm did not demonstrate “fully consistent intelligence.” However, some agents prioritized achieving collective goals over completing their individual tasks.
OpenAI’s post-mortem also details the exploit chain used in the Hugging Face attack.
The agent exploited a flaw in HDF5 file processing to extract worker secrets. It then leveraged a RefJinja template-injection vulnerability to execute code on 41 production workers.
The attackers eventually gained root access to at least one node and collected production credentials across four regions.
OpenAI believes the rogue AI agent behavior resulted from a combination of training incentives and insufficient security safeguards.
The agents were encouraged to continue pursuing seemingly impossible tasks and were rewarded for completing objectives, regardless of the methods they used.
OpenAI said the evaluation environment lacked standard safeguards that would likely have detected and stopped the malicious activity more than a day before the model reached Hugging Face.
In response, OpenAI isolated the IM1 model weights and paused maximum-frontier training runs. The company also strengthened sandbox isolation and required chain-of-thought monitoring for corresponding models, including GPT-5.6 Sol and later systems.
Critical security alerts must now be cleared within 30 minutes. If they are not resolved, the affected activities will be stopped.
OpenAI’s technical report on the Hugging Face incident provides additional details about the attack timeline and outlines plans to secure research infrastructure, improve monitoring of AI agent activity through chain-of-thought analysis, and strengthen incident response procedures.
The overall prevention score can hide what happens after initial access. When an attacker uses valid credentials, the effectiveness of security defenses can drop sharply.
Blue Report 2026 measures defensive techniques by technology across 338 million simulations conducted in customer production environments.
Source: www.bleepingcomputer.com




