How Can We Control and Regulate AI—and Could AI Discussions Shape Future Models?
— Will Douglas Heaven
How can we control, monitor, and regulate AI?
That’s the million-dollar question. Whether you think AI can kill us or not, there’s no denying that AI can cause harm. AI is already linked to real-world harms, including driving people into psychosis and hacking websites. Preventing—or at least mitigating—that damage is difficult for two reasons.
AI systems are becoming more powerful, but remain difficult to understand
First, we have little understanding of how AI works, even as it rapidly becomes more powerful. Although researchers are studying how to monitor and control misbehaving AI agents, current approaches are weak.
OpenAI’s newest agents do not show their work in the same way their predecessors did. However, they can identify whether they are discussing cheating in the “chain of thought,” the workspace where agents plan their actions. Another approach is to use one AI agent to monitor another, but that requires trusting the monitoring agent.
AI companies face conflicts of interest when regulating themselves
The second obstacle is closer to home. Significant conflicts of interest arise when AI companies regulate themselves, while the U.S. government has so far been unable to intervene despite bipartisan support in Congress. For the time being, the executive branch appears to be strongly opposed.
However, if the political winds change, I would support strong transparency regulations. These rules could provide more detailed information the next time an unannounced frontier model launches a cyberattack.
— Grace Huckins
Could online discussions about AI become a self-fulfilling prophecy?
That’s really worrying. Large language models are influenced by what they read. One theory for why chatbots so often discuss and role-play apocalyptic scenarios is that they were trained on millions of pages from science fiction novels and apocalyptic internet forums.
All text created today, including this article, can affect the behavior of future AI models. Very meta.
AI-generated analysis can be influenced by the material it studies
In fact, a team at METR, a third-party organization brought in by OpenAI to understand what happened before the Hugging Face hack, raised a related possibility in its report on the incident. METR used OpenAI’s new model Astra to help analyze a huge number of agent transcripts and behavioral logs.
But when that material is fed into a model, the agent analyzing it may be influenced by text generated by the agent it is studying. There is no such thing as a blank slate anymore.
Thanks to Eric, Pranab, Rafael, Kenneth, George, Chris, Yoon Jae, James, Carl, Nicole, and more for the great questions.
Source: www.technologyreview.com


