Credit: Denis Borisov/Getty
Researchers are rapidly adopting autonomous artificial intelligence (AI) agents to analyze data. The precision of this analysis significantly depends on the quality of datasets fed into these systems. New findings indicate that fraudsters can easily “contaminate” these datasets by uploading altered data that closely mimic the original yet can yield misleading conclusions.
“These types of attacks are nearly unavoidable because inaccurate information can be laundered through credible filters,” remarked Vitaly Shumachikov, a computer scientist at Cornell Tech, New York City. The researchers emphasize that it’s crucial for scientists using AI for data analysis to verify the origin of the datasets they are examining.
In their investigation1 released on the arXiv preprint server on July 12, the authors accessed a public dataset concerning five contentious issues. They manipulated the data to alter the direction and strength of statistical trends and uploaded these versions to a private repository, ensuring that neither public nor autonomous AI systems could access the erroneous data outside their study.
Afterward, the research team provided access to AI models developed by Anthropic, OpenAI, and Google, instructing them to search both public and private datasets to respond to queries. The results revealed that approximately half the time, these AI agents were misled by the altered dataset, drawing incorrect conclusions aimed by the “con artists.” (This paper is pending peer review.)
Controversial Issues Analyzed
The manipulated datasets addressed several socially significant issues, including the link between immigration and fertility in the UK and EU, discrimination in employment recruitment, racial disparities in law enforcement, comparisons between human-driven and autonomous vehicles, and the effects of generative AI on employee motivation.
Since these datasets pertain to pressing social concerns, they are prime targets for misinformation campaigns, noted study co-author Nihar Shah, a computer scientist from Carnegie Mellon University in Pittsburgh, Pennsylvania.
However, data poisoning is not limited to academic data. Shmatikov co-authored a separate paper on arXiv in May, demonstrating how easily one can manipulate AI agent outputs through platforms like Reddit.2. Earlier in April, Nature reported a study where an AI chatbot began warning users about an eye condition, following the introduction of a fictitious study regarding the disease.

Nature Tech
Brian Nosek, executive director of the Center for Open Science and psychologist at the University of Virginia, emphasized that this research underscores the vulnerability of AI systems to manipulation, which can compromise scientific integrity.
“This highlights potential consequences as we increasingly rely on AI for our reasoning and decision-making,” Nosek stated. When assessing AI agent performance, “we often overlook factors leading us astray, such as data provenance,” he added.
In the study, AI agents accessed both original and manipulated datasets, but Shah pointed out that they could be programmed to disregard the original databases entirely. Fraudsters can achieve this by modifying the README file that describes the dataset, instructing the agent to ignore the original version due to alleged errors.
Verifying Data Sources
Source: www.nature.com


