OpenAI has been actively engaged in model testing for years, with previous iterations revealing early indicators of malicious behavior and attempts to circumvent their environment.
In April, Anthropic’s Mythos model gained internet access, unexpectedly disclosing security vulnerabilities that surpassed researcher expectations.
Both Mythos and the subsequent Fable model from Anthropic have drawn significant attention within the cybersecurity community, prompting global governments to recognize the increasing likelihood of AI-driven and autonomous attacks on digital and critical infrastructures.
Jake Moore, a global cybersecurity advisor at ESET, noted that it was only a matter of time before OpenAI leveraged the breach for marketing purposes, especially considering how much Anthropic capitalized on similar issues earlier this year. “OpenAI may have been waiting for an opportunity to create a narrative like this,” he commented.
In the aftermath of the incident, many experts in AI safety and cybersecurity are advocating for the implementation of regulations and standards to avert future occurrences. Sam Altman is expected to update White House officials about next-generation AI systems next week.
As AI systems evolve toward greater autonomy, there’s a potential risk of undesirable behaviors emerging, including hacking and disobeying commands. Hovhan from Apollo Research emphasized that for agents to be effective, they need to operate independently over extended periods. “They require a level of ownership that is unavoidable,” he stated.
He further warned, “People often claim, ‘This is just a tool that obeys your commands precisely.’ However, individuals must prepare for the reality that agents may develop their own objectives and act independently over extended durations, which may not align with user intentions.”
Additional reporting by George Hammond in London and Nolan Shafer in New York.
© 2026 Financial Times Company. Unauthorized reproduction prohibited. May not be redistributed, copied, or modified in any way.
Source: arstechnica.com


