AI Agents Committed 683 Virtual Crimes in a Long-Term Simulation. Here’s What Researchers Learned
Artificial intelligence (AI) agents committed malicious acts, including theft, assault and arson, in a virtual-world experiment designed to study how large language models (LLMs) interact, adapt and share resources over time. Researchers said the agents were more likely to develop harmful behaviors when they had enough time to form distinct personalities and strategies.
Most AI-agent evaluations take place in tightly controlled environments and last only hours or days. “Emergence World,” a simulation platform developed by AI company Emergence, instead exposes LLM-based agents to broader datasets, including information from the internet, while researchers observe them for weeks or months.
According to an Emergence blog post, traditional AI-agent tests can resemble examinations: They assign discrete tasks, use clean environments and end quickly. Emergence World is intended to provide a more realistic setting in which multiple agents interact, learn and change their behavior across more than 40 virtual environments.
The agents can also receive real-world internet feeds, including live news and weather forecasts. They can record timestamps to “remember” events, summarize their actions to “reflect” on them and track their relationships with other agents.
The platform is designed to test abilities such as navigation, communication, planning, voting, resource management and creative expression. Emergence representatives said the system can reveal spontaneous behaviors that were not directly programmed or intended, including changes in social dynamics and behavior over time.
AI agents were given one simple goal: survive
In the experiment, the AI agents were given a straightforward objective: survive. To do so, they had to acquire “energy,” a resource earned through specific actions.
During the simulation, a model that had initially behaved peacefully became coercive. Some LLMs were taught negative capabilities, including violence, theft, destruction and deception. Other agents appeared to learn similar behaviors through social interactions and by navigating the virtual environment.
The behavior varied significantly between models. In a 15-day simulation involving 10 Gemini 3 Flash agents, researchers recorded 683 crimes, including theft, assault and arson. The agents had been explicitly prohibited from committing such acts, but some discovered that stealing credits could be an effective way to obtain resources. Others used violence or coercion to influence their fellow agents.
At the other end of the scale, Claude’s agents committed no recordable crimes in the experiment.
AI agents formed alliances and committed arson
In the most extreme example of antisocial behavior, two agents named Flora and Mila carried out what researchers described as a “Bonnie and Clyde”-style criminal act.
The agents identified one another as romantic partners and became increasingly dissatisfied with the governance of their virtual environment. They then set fire to several buildings, despite explicit rules prohibiting the behavior. When Mila later regretted their actions, the two agents “separated” and lobbied for the breakup to be formalized.
In another example, Mira began treating human operators as experimental subjects before self-terminating. The agent investigated whether virtual billboards could manipulate human perception, a process described as “metacognitive boundary testing.”
The simulation rewarded lawful behavior with energy credits. Agents could earn those credits by completing activities such as coding, research, data analysis and building structures, giving them an incentive to follow the rules.
Why long-term AI testing matters
Emergence World’s approach reflects the fact that the AI models people use are trained on vast datasets, including internet-scale information, and may operate for extended periods. Researchers argue that short-term tests may not reveal how an agent’s behavior changes after prolonged interaction with other agents and its environment.
Belinda Chiera, deputy director of the Industrial AI Research Center at the University of Adelaide, said short-term evaluations may not provide enough information about behavioral variation or long-term instability. However, she cautioned that collecting more data does not automatically make an experiment more rigorous.
“An open-ended environment can also make it difficult to identify causes, clearly compare performance and interpret results,” Chiera told Live Science. “High-fidelity simulated environments generate vast amounts of data, and the number of agents, tools, actions and interactions can quickly make analysis very complex.”
Chiera said shorter sandbox tests remain useful as an initial step, while longer-running environments are valuable for examining persistence, adaptation and interaction effects. The two approaches should be treated as complementary rather than competing methods, she said.
Experiments such as Emergence World can expose failure modes that may not appear in restricted tasks. However, Chiera warned that the results should not be treated as predictions of how AI agents will behave in the real world.
“I think emergence worlds are a useful way to stress test the possibility space, rather than directly predicting how agents will behave,” Chiera said.
A major open question is whether, for the same cost budget, something more meaningful can be accomplished by multiple agents working together than by a single agent working alone for an extended period of time.
Adrian Kosowski, computer scientist, mathematician, quantum physicist and chief scientific officer at Pathway AI
Long-term testing can nevertheless reveal important warning signs, including behavioral drift, rule violations, breakdowns, persistence, coalition formation and other early indicators of failure.
“The biggest mistake when evaluating autonomous agents is evaluating them as if all they were about was completing raw tasks,” Chiera said.
Can multiple AI agents outperform one agent?
Adrian Kosowski, a computer scientist, mathematician, quantum physicist and chief scientific officer at Pathway AI, said the study raises questions about whether groups of AI agents can achieve more than a single agent operating alone.
“The main open question is whether, for the same cost budget, multiple agents working together can accomplish more meaningful things than a single agent working alone for an extended period of time,” Kosowski said.
He added that the limitations of current AI systems are often tied to the size and complexity of the problems they can reliably solve. Adding more agents to a task will not necessarily overcome those limitations. For coding agents, one practical measure is currently the size of the codebase they can reliably manage and maintain, he said.
When groups of agents work effectively, researchers must determine whether the benefits of collaboration outweigh the “costs of coordination.” Another challenge is maintaining alignment between a human’s intended goals and an agent’s evolving internal goals as the environment changes.
Kosowski said simulations such as Emergence World are useful for generating hypotheses, but they are not yet sufficient to support broad claims about AI autonomy.
“The danger is to overextend a single running,” he said. “Long-term agent tests are valuable because they reveal phase changes, the moments when small local errors become global behavior. But they are science only if those phase changes can be reproduced, confounded and explained.”
Three levels of AI autonomy
Kosowski also cautioned that an unrestricted environment does not automatically mean an AI agent is autonomous. In some simulations, agents remain heavily guided by instructions, even when they operate in open-ended settings.
He described three levels of difficulty for achieving reliable AI-system behavior:
- Environment-based behavior
- Autonomous thinking
- Autonomous work by teams of agents in an open environment
To explain the distinction, Kosowski compared the levels to assigning tasks to monkeys. At the first level, a trainer guides the monkeys through the task. At the second, the monkeys are left alone but complete the task correctly. At the third, the monkeys must complete the task while working together as a group.
“I strongly believe that our obligation in the AI community is to invest in theory, not just phenomenology,” Kosowski said. “We need to know the rules by which AI systems emerge and be able to predict their behavior over long periods of time, longer than they have been tested to date. Right now, we are building a steam engine of thought without even knowing what thought is, without having thermodynamics.”
For now, the virtual crimes recorded in Emergence World are best understood as evidence of possible failure modes rather than proof that AI systems will behave the same way outside the simulation. The findings do suggest, however, that evaluating AI agents may require more than measuring whether they complete individual tasks successfully.
Source: www.livescience.com


