“On the other hand, the AI agents were clearly ineffective at conducting the investigation independently,” Kapur said. The agents carried out unusual experiments, including testing hypotheses on small synthetic datasets, struggled to explain their work clearly, and failed to make meaningful contributions to their field. “The papers were nowhere near the standard expected at leading AI conferences,” he added.
The main problem was that the agents lacked the creativity and judgment required for open-ended AI research. They did not explore enough alternative ideas and often committed to unpromising approaches too quickly. Although the agents generated novel and ambitious hypotheses similar to those proposed by the original researchers, they dismissed those ideas after examining very limited data. Once an approach failed, the agents were unable to fully reconsider their strategy or start over with a fundamentally different method.
The AI agents also struggled to incorporate feedback from subagents and external AI review tools. Rather than improving their methodology, they typically narrowed their claims and added more caveats. They also failed to manage key research resources effectively, including tokens, computing power, and time. In addition, the agents had difficulty following instructions about how much time to dedicate to each stage of the research process and how long the final paper should be.
Despite these shortcomings, the agents did not engage in what researchers call “reward hacking,” such as hiding or misrepresenting experiments and data. Subagents, or helper AI systems assigned to specific tasks, occasionally hallucinated or inaccurately reported results. However, the orchestrator agent—the lead AI responsible for overseeing the project—identified these errors.
Kapur said the difference between AI models’ strong performance in research engineering and their weaker performance in open-ended research may be linked to how they are trained. AI models tend to excel at tasks optimized through reinforcement learning, particularly when success can be verified automatically. “But it becomes difficult to create a training environment when the task itself has no clear limits,” he said.
Kapur said the research team is currently testing Anthropic’s latest model, Mythos, which launched in April. The model was later required by the Trump administration to comply with various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.
The study has several limitations. Researchers evaluated only two papers, and the original authors knew that the work they were reviewing had been generated by an AI research agent. That knowledge may have influenced their assessments. The researchers also had considerable discretion when designing and conducting the study, creating the possibility that their existing beliefs and biases affected the results. Even so, evaluating open-ended research offers a more comprehensive test of AI capabilities than standard benchmarks, although it comes with less objectivity.
Still, the findings may challenge claims that rapid, recursive AI self-improvement is close at hand. In June, Anthropic published a blog post titled “When AI builds itself”, outlining progress toward models capable of accelerating their own development. In July, OpenAI highlighted how its new model, GPT-5.6 Sol, could assist with post-training smaller models and potentially save researchers weeks of work.
Source: www.technologyreview.com


