Research

OpenAI, Google, and Anthropic Models Cite Same Papers

A new study reveals that major language models from OpenAI, Google, and Anthropic repeatedly cite the same narrow set of papers, threatening to create a scientific citation monoculture.

Unite.AI23 hrs agoResearch
Image: Unite.AI

A collaborative study by researchers from the University of Texas at Austin, Stevens Institute of Technology, Washington University at St Louis, Rice University, and the University of Notre Dame has found that major AI models suffer from a citation monoculture. The paper, titled When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models, warns that these systems repeatedly favor a narrow group of scientific papers, potentially drowning out novel research.

The researchers evaluated 11 models from three major AI vendors. The tested systems included OpenAI's GPT-5, GPT-5 mini, GPT-4.1, and GPT-4.1 mini; Google's Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 3.1 Pro, and Gemini 3.1 Flash-Lite; and Anthropic's Claude Opus 4.8, Claude Sonnet 4.6, and Claude Haiku 4.5. The team gathered a benchmark of 120 real knowledge-distillation papers published on arXiv between 2015 and 2022, each having between 50 and 500 citations.

To isolate bias, the researchers stripped the papers of citation counts, venues, and author names, and randomized the publication years. When shown random sets of 30 papers and restricted to citing a maximum of 10, all 11 models consistently converged on the same small subset of papers. In contrast, eight human experts given the same blinded papers did not show this concentration bias. For GPT-5 mini and GPT-4.1 mini, about 90 percent of this preference variation was driven by semantic content rather than metadata or list order. Rewriting titles and abstracts to preserve meaning still yielded a high correlation of 0.95 to 0.99 across the models.

In an extended experiment spanning 11 rounds, the researchers added 120 AI-generated papers after each round, eventually expanding the pool to 1,440 papers. As the proportion of AI-written literature grew, the models increasingly concentrated their citations on an even smaller, shrinking number of the original human-authored papers. The authors suggest that simply mixing models or giving papers equal exposure will not solve this systemic bias, meaning researchers may need to actively boost neglected papers to maintain scientific diversity.

This is our own summary of reporting by Unite.AI

More in Research