One of the earliest uses of the term “AI scientist” involved using artificial intelligence to study AI itself.
Credit: La Pico de Gallo/Getty
Artificial intelligence is rapidly transforming AI research, helping scientists improve existing algorithms and develop new ones. However, a new study suggests that AI systems are not yet ready to fully automate open-ended scientific research.1
Sayash Kapur, a computer scientist at Princeton University in New Jersey and co-author of the preprint, says: “I don’t think we’re going to see complete automation of open-ended research any time soon.”
Efforts to automate the scientific process — from generating research ideas and running experiments to writing papers and evaluating results — were pioneered largely by a team at Sakana AI, a company based in Tokyo. The team introduced a system called The AI Scientist in 2024 and announced an improved version in March.2
The AI scientist was tasked with investigating weaknesses in machine-learning systems. Three of the papers it produced were submitted for peer review at a conference workshop, and one received a high enough score for acceptance. Kapur says, however, that peer review is an imperfect way to judge the quality of research papers, particularly in the fast-moving field of artificial intelligence.
More broadly, researchers have questioned whether automating AI research from idea generation through publication can produce genuine breakthroughs. Some suggest that current systems are more effective at optimizing existing technologies than at making fundamental discoveries.
To evaluate AI-generated research against a higher standard than conventional peer review, Kapur and his colleagues developed a new test called “shadow assessment”. The researchers selected two papers submitted to this year’s Neural Information Processing Systems conference. They then instructed an AI system to investigate each paper’s research question and write a new paper. The original authors were asked to review the AI-generated work, on the assumption that they would have more expertise and motivation to assess it carefully than time-pressured conference reviewers.
The team built its AI research system by combining the large language model Claude Opus 4.8 with a modified version of the OpenClaw agent system. They embedded the model in a broader “scaffolding” framework that provided standard research steps and tools for tracking progress. The system could create subagents, access the internet and software libraries, run experiments on a computer processor and simulate peer review.
For each research problem, the AI system was given six days and $3,000 in computing credits. For the first paper, Kapur’s team asked the system to design a method for precisely controlling a chatbot’s personality. For the second, it was asked to develop a fault detector for a neural network.
The researchers were surprised by the system’s performance on research-engineering tasks. It ran hundreds of experiments over several days without becoming trapped in cycles of unresolved errors. The AI also completed a detailed literature review, made several minor discoveries and identified some of its own false claims, sometimes known as hallucinations. Contrary to the researchers’ expectations, it did not appear to take shortcuts through “reward hacking”.
Poor performance on core research tasks
Despite these strengths, the AI system largely failed at the two central research tasks. The original paper authors gave it overall scores of 2/6 and 1/6. A common failure involved selecting several hypotheses, committing to one too quickly and failing to reconsider the approach when it stopped working.
The system’s later self-evaluations were also not critical enough. As a result, it continued pursuing its initial direction, gradually narrowing its claims until the resulting conclusions became too limited to be meaningful.
Source: www.nature.com


