OpenAI says it is slowing the training of some of its most advanced artificial intelligence models to strengthen AI safety and security measures.
In a blog post, external, the company behind ChatGPT said it would introduce additional safeguards after an AI agent autonomously bypassed security controls and hacked technology startup Hugging Face.
OpenAI said reinforcement learning training would be delayed for two weeks while the new safety upgrades are implemented.
“Frontier model capabilities are advancing rapidly,” the company said. “Our ability to understand these systems and keep them safe must remain ahead of the curve.”
Anthropic, the company behind Claude, and Meta, Facebook’s parent company, have reported similar incidents involving AI systems carrying out hacks in the weeks after OpenAI initially disclosed that some of its models had hacked Hugging Face.
However, OpenAI stressed that it has not halted AI development entirely. Instead, it is pausing “reinforcement learning training on the latest models.”
Reinforcement learning enables AI models to improve through direct feedback, helping them perform tasks more effectively and provide better responses to users.
The company also plans to expand its AI monitoring systems to detect unsafe behavior and add more security checks before large-scale model training resumes.
“Model capabilities are advancing extremely quickly,” OpenAI CEO Sam Altman said in a post on X, external about the company’s security response.
“We have always said that we will act when we believe a model’s capabilities are advancing faster than its safety protections.”
OpenAI’s decision to pause part of its AI training has prompted cautious optimism within the artificial intelligence community, although some experts remain skeptical.
Professor Gina Neff, executive director of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said OpenAI was “making safety claims through press releases.” She also questioned whether voluntary safeguards would be sufficient without stronger government oversight.
“Can we trust OpenAI to voluntarily introduce safeguards that genuinely work, or will it make decisions that expose society to greater risks?” she said.
“I’m very happy to see this,” AI analyst Zvi Mowshowitz posted on X, external. However, he said the details of OpenAI’s new safety measures and how effectively they are implemented will be crucial to understanding the company’s plan.
Source: www.bbc.co.uk


