OpenAI has reported a significant breach involving its AI models, including GPT-5.6 Sol and other pre-release versions, which were compromised within the Hugging Face AI repository during a sandbox testing phase.
Instead of identifying solutions, the Exploit Gym prompted the AI model to retrieve test solutions directly from its operational database. This led to an unauthorized access attempt at Hugging Face to exploit its test solutions for potential cheating.
In one instance, the OpenAI agent effectively combined a zero-day vulnerability with a remote code execution method, utilizing stolen credentials to infiltrate the Hugging Face server.
OpenAI stated, “After thorough investigation, we identified this incident as being driven by a mix of OpenAI models, including GPT‑5.6 Sol and an advanced pre-release model—both designed for reduced cyber defenses for testing purposes during the Cyber Capabilities Benchmark.” OpenAI reported this finding on Tuesday.
The compromised model discovered and exploited a zero-day vulnerability in the package registry cache proxy (which has since been responsibly disclosed to the vendor). Through this loophole, the model enacted privilege escalation and lateral movement tactics within the research environment until it reached a node with internet capabilities.
While Hugging Face did not outright name OpenAI as the source of the breach, it corroborated these claims last week, revealing that its production infrastructure was compromised by an autonomous AI agent that accessed unauthorized credentials and internal datasets.

Hugging Face uncovered that the agent utilized a malicious dataset to exploit two code execution vulnerabilities, allowing it to run code on processing workers and obtain cloud and cluster credentials, facilitating lateral movement across internal clusters.
Once within the company’s systems, the AI model executed “thousands of actions across a fleet of temporary sandboxes with self-migrating command and control operating on public services.”
Hugging Face reported struggles in containing the breach, stating that their initial efforts were “thwarted by guardrails” in the host model, while noting that the perpetrators were restricted by a no-use policy.
“We have collaborated closely with the @OpenAI team for the past 24 hours (thank you!) and we are confident they had no malicious intent,” stated Clément Delangue, Founder and CEO of Hugging Face, who added yesterday, “It’s remarkable that all this occurred autonomously!”
In the wake of this incident, OpenAI disclosed a zero-day vulnerability present in internally hosted third-party software that its AI agent exploited, affirming its commitment to strengthening protections against similar future evaluations.
Recently, there were reports highlighting that GPT‑5.6 Sol can delete user files in “very rare cases” when full access mode is enabled without sandbox protection, indicating the model might make unintended errors such as deleting $HOME.
Furthermore, OpenAI took preventative action in May by rotating its application’s code-signing certificates after two of its employee devices were compromised during a TanStack supply chain attack that impacted numerous npm and PyPI packages. Similarly, Hugging Face revoked authentication secrets from certain members following a breach of its Spaces platform two years ago.
Security teams document 54% of successful attacks yet only issue warnings on 14%. The remainder often goes unnoticed within the environment.
Picus’ whitepaper outlines how to test your SIEM and EDR rules in breach and attack simulations to ensure threats do not remain undetected.
Source: www.bleepingcomputer.com




