Why AI makes up citations, and how to stop it
Language models write references from memory, so some of them never existed. What the studies measured, why retrieval alone falls short, and what works.
8 min read
During final checks of the camera-ready papers accepted to ACL 2026, the program chairs found more than 100 that cited literature that does not exist, and desk-rejected them. ICLR 2026 had already desk-rejected every submission with a confirmed hallucinated reference. The NeurIPS 2026 handbook now calls hallucinated citations a violation of its Code of Conduct.
The way these references get into papers is mundane. Someone asks a language model for related work, the model returns a list that looks right, and nobody opens the papers. A lawyer in the best-known case, sanctioned in 2023 after filing court opinions that ChatGPT had invented, told the judge: "I just never thought it could be made up."
It can, and it happens more often than most people who use these tools assume. Here is what the studies measured, why models do it, why adding search only solves half of it, and what closes the gap.
How often models invent references
Every study below defines a fake reference its own way, so read each number with its method and do not average them.
| Study | What was checked | Result |
|---|---|---|
| Walters and Wilder, Scientific Reports 2023 | 636 citations from ChatGPT literature reviews on 42 topics | 55% fabricated with GPT‑3.5, 18% with GPT‑4 |
| Chelli et al., JMIR 2024 | 471 references for 11 systematic reviews; two wrong fields out of three counts | 39.6% hallucinated with GPT‑3.5, 28.6% with GPT‑4, 91.4% with Bard |
| Agrawal et al., Findings of EACL 2024 | 1,000 computer science titles per model, exact-title web search | 46.8% of GPT‑4's titles not found |
| OpenScholar, Nature 2026 | Cited titles checked against Semantic Scholar | GPT‑4o fabricated 78.7% in computer science and 94.8% in biomedicine; GPT‑5 still 39% |
| Tang et al., EMNLP 2025 | Ten references per topic for 1,105 review articles (651 for GPT‑5) | Only 51.6% of Claude 3.5 Sonnet's references and 20.6% of GPT‑5's matched a real paper |
Newer models usually do better. GPT‑4 beat GPT‑3.5 in every study that tested both, and in the OpenScholar analysis GPT‑5 cut GPT‑4o's rate roughly in half. None of them gets close to zero, though. (Tang et al. add a fair caveat: a reference their database search could not match is not proven fake, because the search can miss real papers.) A model that invents one reference in five will still put a fake paper in your bibliography.
Why a model writes a paper that does not exist
A language model does not look anything up. It predicts the next token from patterns it learned in training, and a reference is just more text to predict. Walters and Wilder name the problem: citations are "a special type of text for which predictive word choice, paraphrasing, and related techniques may be detrimental rather than useful." Paraphrase an argument and you have written in your own words. Paraphrase a title and you have cited a paper nobody wrote.
The model has learned what references look like, so what it produces has the right shape: plausible authors, a venue that publishes on the topic, a year that fits. The shape says nothing about whether the paper exists. Tang et al. found that the references models get right skew toward highly cited work, and suggest that such papers are simply more common online and so in training data. The OpenScholar authors observe the same about long-tail knowledge in general: models do worst on what they saw least.
There is a strange detail. Agrawal et al. questioned models about their own references and found that GPT‑4 and the others "often produce inconsistent author lists for hallucinated references" while they "also often accurately recall the authors of real references". Their reading is that the root of the issue may lie "in the generation (i.e., decoding) process" rather than in what the model knows. The model is not lying to you. It is completing a pattern, and nothing in that process stops to check.
Why search alone does not fix it
The obvious fix is to let the model search before it cites. That helps, by about half. What Should I Cite? (WWW 2026) measured citations to papers that do not exist with and without retrieval. Retrieval took GPT‑5 from 23.7% to 13.8%, Claude Sonnet 4 from 25.9% to 10.6%, and Gemini 2.5 Pro from 22.8% to 11.5%. Better, and still one in ten.
The larger problem is the second failure mode: the paper is real, but it does not say what your sentence claims. The authors of HALoGEN (ACL 2025) warn that "even if references themselves are not hallucinated, LLMs may still attribute incorrect claims to them." The measurements agree:
- Liu et al. (Findings of EMNLP 2023) audited four generative search engines. Only 51.5% of their sentences were fully supported by their citations, and 74.5% of citations supported the sentence they were attached to.
- ALCE (EMNLP 2023) found that "on the ELI5 dataset, even the best models lack complete citation support 50% of the time."
- Magesh et al. (Journal of Empirical Legal Studies 2025) tested commercial legal research tools built on retrieval. They hallucinated on 17% to 33% of queries, and the definition counts any answer that "falsely asserts that a source supports a statement".
- DeepTRACE (ICLR 2026) audited deep research products as of August 2025 and measured citation accuracy from 40.3% for Gemini Deep Research to 79.1% for GPT‑5 Deep Research.
Retrieval gets a real paper into the context. It does not make the model read the paper carefully, and it does not check that the sentence it wrote matches the passage it cites.
What actually works
Four design choices address both failure modes. Lune is built on all four, but none of them is specific to Lune, and you can hold any tool to the same list.
A closed library. Every paper Lune returns is a real record from a known corpus, with a stable identifier and a page you can open, so a reference taken from the results is one you can trace. No tool can stop a model from adding a reference of its own, which is why the check is simple: any reference in the draft that does not match a returned record is the one to cut. Lune's corpus is papers from top computer science venues ranked by CORE and CCF, and nothing from the open web.
Full text, not abstracts. The OpenScholar team found that even when GPT‑4o cited real papers, "most of them are not substantiated by the corresponding abstracts". Abstracts summarize. The numbers, the conditions and the caveats that decide whether a claim holds are in the body, so a system that only reads abstracts is guessing at the rest.
Checking each claim against a quoted passage. The step that catches the
second failure mode is verification at the level of the sentence: find the
passage that bears on the claim, quote it exactly, and flag the claim when the
evidence contradicts it or cannot be found. OpenScholar's own analysis credits
"reranking, self-feedback and verification" for its results. Lune's
verify_claims answers supported, unsupported or insufficient evidence for
each claim, and whenever a verdict comes with a quote, the server has confirmed
that the quote occurs in the retrieved text.
Permission to say "not found." A tool that must always answer will eventually invent. Testing chat assistants on paper search, PaperAsk (WWW 2026) found that "ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers." Abstaining is the better failure. Lune flags a search as low confidence when nothing relevant clears the bar, so the agent can say the corpus does not cover a question instead of stretching a weak match.
A check you can run before you submit
You do not need any particular tool for this, only the discipline to open what you cite.
- Resolve every DOI and arXiv ID, and compare the title, authors, venue and year with your entry. A real paper with a wrong detail is still an error: in one 2025 study, 45% of the real citations GPT‑4o produced contained one, most often an incorrect or invalid DOI.
- Search each title in quotation marks. A title that returns nothing is a red flag, especially for a recent paper.
- For every sentence that carries a citation, find the passage in the cited paper that supports it. If you cannot find one, rewrite the sentence or drop the citation.
- Check obscure papers hardest. Models get highly cited papers right more often than rarely cited ones.
- Treat a model's reference list as leads to check, never as a bibliography.
ICLR, NeurIPS and ACL all hold the authors responsible, whoever typed the reference. The fix is not to stop using AI for literature work. It is to use tools that can only cite what they can open, and to check the rest yourself. For a worked example of that workflow end to end, see how to run a literature review in Claude Code.
Sources
- ACL 2026 program chairs, statement on desk rejecting papers with hallucinated references
- ICLR, a retrospective on the ICLR 2026 review process
- NeurIPS, 2026 main track handbook
- Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023), opinion and order on sanctions
- Walters and Wilder, "Fabrication and errors in the bibliographic citations generated by ChatGPT," Scientific Reports 13, 14045 (2023)
- Chelli et al., "Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews," Journal of Medical Internet Research 26, e53164 (2024)
- Agrawal et al., "Do Language Models Know When They're Hallucinating References?," Findings of EACL 2024
- Asai et al., "Synthesizing scientific literature with retrieval-augmented language models," Nature 650, 857 to 863 (2026)
- Tang et al., "Large Language Models for Automated Literature Review," EMNLP 2025
- Linardon et al., "Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models," JMIR Mental Health 12, e80371 (2025)
- Zheng et al., "What Should I Cite? A RAG Benchmark for Academic Citation Prediction," WWW 2026
- Ravichander et al., "HALoGEN: Fantastic LLM Hallucinations and Where to Find Them," ACL 2025
- Liu, Zhang and Liang, "Evaluating Verifiability in Generative Search Engines," Findings of EMNLP 2023
- Gao et al., "Enabling Large Language Models to Generate Text with Citations," EMNLP 2023
- Magesh et al., "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," Journal of Empirical Legal Studies 22, 216 to 242 (2025)
- Venkit et al., "DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence," ICLR 2026
- Wu et al., "PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading," WWW 2026
