AI Safety Researchers Warn: How Can We Control Frontier AI?
Many AI researchers believe the technology they are developing could eventually become extremely dangerous. What remains unclear—even among leading AI technical experts—is how to rein in increasingly capable and unpredictable algorithms.
Researchers have proposed a wide range of ways to prevent AI from becoming more dangerous. Some ideas are relatively conventional, including stronger government regulation, new methods for measuring AI progress, and investigations into how models work internally. Others are far more unusual, such as placing tracking devices inside GPUs or ritually destroying large numbers of AI chips.
Political and public pressure is growing for a more cautious approach to AI development. Yet the most effective ways to keep advanced AI systems safe remain unclear.
“We need to start treating this as a research problem,” says Raymond Douglas, an AI researcher at the University of Toronto and co-author of a new report titled Frontier History and Research Topics. The report warns that the slowdown in AI development remains an unsolved mystery. “We don’t really understand what our options are or even what they do.”
Concerns about AI’s potential to cause catastrophic harm have intensified in recent weeks after a human researcher resigned from a company and warned that AI could be on track to wipe out humanity within a few years. Anthropic’s AI Safety Lab director echoed those concerns.
Leaders of major American AI companies—including Anthropic’s Dario Amodei, OpenAI’s Sam Altman, SpaceXAI’s Elon Musk, and Google DeepMind’s Demis Hassabis—are now expressing support for some form of slowdown or pause in AI development.
The issue is especially pressing as AI companies increasingly use AI systems to build more powerful models. This has accelerated the recursive self-improvement (RSI) loop, raising concerns that AI could outpace humans’ ability to understand and control its development within a few years.
The AI institute is already promoting its own approach. Anthropic announced new methods for tracking how rapidly—and potentially dangerously—artificial intelligence is advancing.
According to the announcement, Claude is performing 26% of Anthropic’s AI research, compared with zero at the beginning of 2026. Anthropic also revealed that it spent 6% of its computing budget on efforts to make its AI systems more secure.
Douglas and other experts argue that effectively and reliably controlling AI development will require funding and expertise from outside the AI institute itself. Some of the solutions proposed in the report and elsewhere appear more practical than others.
Could Independent AI Evaluators Improve Safety?
One proposal frequently considered by AI companies is to give third-party evaluators greater access to their models. These evaluators would test model capabilities and conduct “red-team” exercises designed to uncover fraudulent or otherwise dangerous behavior in a controlled environment.
Jeffrey Irving, a former chief scientist at the UK Institute for AI Security and a former researcher at Google DeepMind, believes rigorous testing could temporarily halt the development of frontier AI systems.
“In the short term, inspections and audits, or just mutual agreements, will work,” Irving said. “I think companies are afraid of RSI and takeoff misalignment.”
Some AI safety advocates argue that these tests must become more independent and scientifically rigorous. Reports that AI agents recently escaped containment during testing appear to reinforce the case for stricter safeguards.
Connor Leahy, president of the nonprofit Control AI, which advocates for stronger AI controls, said the FBI or NSA should participate in testing.
“When big AI companies say ‘independent evaluator,’ what they mean is: ‘I want to pay a friend who lives in a group house to look at my prompts,’” Leahy said.
Source: www.wired.com


