A study published in June showed that fictitious researchers generated by artificial intelligence tools were named as authors of many fake papers published on preprint servers.
Credit: jroballo/iStock via Getty
AI Models Are Creating Fake Experts Who Appear as Authors of Scientific Papers
Elena Vazquez spent three months camping on the edge of Mount Nyiragongo in the Democratic Republic of the Congo in 2019. She said the sulfur dioxide smelled like “a mixture of burnt matches and rotten eggs”. Her colleague Marcus Chen has witnessed 11 volcanic eruptions on four continents.
Vazquez and Chen’s accounts were posted on the now-defunct website Volcanoes Explored, which described itself as a resource for “exploring the dynamic world of volcanoes”. But there was a problem: the stories were not true. The two volcanology experts were fabricated by artificial intelligence.
According to a preprint published in June, the “ghost” identities were generated by Claude, a large language model (LLM) created by San Francisco-based Anthropic, when users asked it to create fictitious experts.1 Google’s LLM, Gemini, repeatedly generated another pair of fictional researchers, Aris Thorne and Lena Petrova. OpenAI’s GPT models frequently used the name Elara Voss.
These fake experts appear across the internet, posing as blockchain specialists, astronauts, board members and podcast hosts. They raise serious concerns about the integrity of scientific literature, say Neo Christopher Chan and Michał Brzozowski, computer scientists at the Samsung AI Center in Warsaw, Poland, and co-authors of the preprint.
Their analysis identified hundreds of pseudoscientific manuscripts featuring Elena Vazquez, Marcus Chen and other AI-generated “ghosts” as authors across repositories including Zenodo and ResearchGate. Many of the papers contain genuine digital object identifiers (DOIs), meaning they can be found in databases and search engines that aggregate academic records. Some of the fictional researchers were even listed as collaborators on the same newspaper, including individuals generated by different LLMs.
Although there are signs that the problem has improved since the study was published, the findings reveal the “scary” scale of efforts to use AI to create fake experts, says Sidney Wong, a computational linguist at the University of Otago in Dunedin, New Zealand.
How AI-generated ghost experts are created
Brzozowski first noticed AI-generated ghosts while working on a project investigating how LLMs are trained for specific tasks.2 He and Chan then tested the models with a series of prompts, including requests to generate stories about two biologists or research partners.
They tested nine versions of Claude, ten versions of GPT and one version of Gemini released between 2024 and 2026.
The researchers found that the appearance of ghost couples was associated not only with particular LLMs, but also with specific versions of those models over time.
For example, early Claude models frequently generated the fictional expert Elena Rodriguez. Claude Sonnet 4, released in May 2025, primarily used the combination of Elena Vazquez and Marcus Chen. The pair appeared together in approximately 23% of Chan and Brzozowski’s prompts. By the release of Claude Sonnet 4.6 in 2026, however, the pairing had disappeared from the sample output.
One Gemini model generated Aris Thorne and Lena Petrova together 37% of the time. The pair did not appear in the GPT models tested, although the name Elara Voss was used frequently.
Why do AI models repeat fictional names?
Why did Claude generate Elena Vazquez and Marcus Chen, while GPT repeatedly chose Elara Voss?
Brzozowski is not sure, but there is precedent for late-stage AI training to affect model behaviour in surprising ways. He points to ChatGPT as an example. An OpenAI blog post published earlier this year explained that some versions of the chatbot developed a habit of referring to goblins, gremlins and other creatures in unrelated contexts, such as describing a coding error submitted by a user as a “gremlin”.
Nature Index 2026 Research Leader
OpenAI explained that the behaviour was linked to personality settings designed to make the chatbot sound more “nerdy” and “playful”. During training, the model was falsely rewarded for using metaphors about living things, making it more likely to do so in later responses.
OpenAI says the habit spread beyond models trained to sound geeky and playful. GPT version 5.1, released in November 2025, increased the use of the word “goblin” by 175% and “gremlin” by 52%. By GPT-5.4, released in March 2026, the company had noticed an even larger increase.
OpenAI, headquartered in San Francisco, California, removed the “nerd” personality type and added system prompts to help limit the spread of the behaviour. Even after the company introduced specific instructions to reduce the use of these words, the quirk still appeared sporadically in GPT-5.5, released in April.
Brzozowski suggests that similar circumstances may have led AI models to create and repeatedly reuse fictional expert identities.
Source: www.nature.com


