# Why AI makes up citations, and how to stop it

Language models write references from memory, so some of them never existed. What the studies measured, why retrieval alone falls short, and what works.

Published 2026-10-01

During final checks of the camera-ready papers accepted to ACL 2026, the program
chairs found more than 100 that cited literature that does not exist, and
desk-rejected them. ICLR 2026 had already desk-rejected every submission with a
confirmed hallucinated reference. The NeurIPS 2026 handbook now calls
hallucinated citations a violation of its Code of Conduct.

The way these references get into papers is mundane. Someone asks a language
model for related work, the model returns a list that looks right, and nobody
opens the papers. A lawyer in the best-known case, sanctioned in 2023 after
filing court opinions that ChatGPT had invented, told the judge: "I just never
thought it could be made up."

It can, and it happens more often than most people who use these tools assume.
Here is what the studies measured, why models do it, why adding search only
solves half of it, and what closes the gap.

## How often models invent references

Every study below defines a fake reference its own way, so read each number
with its method and do not average them.

| Study                                                                                             | What was checked                                                               | Result                                                                                 |
| ------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------- |
| [Walters and Wilder](https://www.nature.com/articles/s41598-023-41032-5), Scientific Reports 2023 | 636 citations from ChatGPT literature reviews on 42 topics                     | 55% fabricated with GPT‑3.5, 18% with GPT‑4                                            |
| [Chelli et al.](https://www.jmir.org/2024/1/e53164), JMIR 2024                                    | 471 references for 11 systematic reviews; two wrong fields out of three counts | 39.6% hallucinated with GPT‑3.5, 28.6% with GPT‑4, 91.4% with Bard                     |
| [Agrawal et al.](https://aclanthology.org/2024.findings-eacl.62/), Findings of EACL 2024          | 1,000 computer science titles per model, exact-title web search                | 46.8% of GPT‑4's titles not found                                                      |
| [OpenScholar](https://www.nature.com/articles/s41586-025-10072-4), Nature 2026                    | Cited titles checked against Semantic Scholar                                  | GPT‑4o fabricated 78.7% in computer science and 94.8% in biomedicine; GPT‑5 still 39%  |
| [Tang et al.](https://luneresearch.com/papers/e84b5dc7-47b0-4cbe-8f6d-a8da1b37040a), EMNLP 2025                           | Ten references per topic for 1,105 review articles (651 for GPT‑5)             | Only 51.6% of Claude 3.5 Sonnet's references and 20.6% of GPT‑5's matched a real paper |

Newer models usually do better. GPT‑4 beat GPT‑3.5 in every study that tested
both, and in the OpenScholar analysis GPT‑5 cut GPT‑4o's rate roughly in half.
None of them gets close to zero, though. (Tang et al. add a fair caveat: a
reference their database search could not match is not proven fake, because
the search can miss real papers.) A model that invents one reference in five
will still put a fake paper in your bibliography.

## Why a model writes a paper that does not exist

A language model does not look anything up. It predicts the next token from
patterns it learned in training, and a reference is just more text to predict.
Walters and Wilder name the problem: citations are "a special type of text for
which predictive word choice, paraphrasing, and related techniques may be
detrimental rather than useful." Paraphrase an argument and you have written in
your own words. Paraphrase a title and you have cited a paper nobody wrote.

The model has learned what references look like, so what it produces has the
right shape: plausible authors, a venue that publishes on the topic, a year
that fits. The shape says nothing about whether the paper exists. Tang et al.
found that the references models get right skew toward highly cited work, and
suggest that such papers are simply more common online and so in training data.
The OpenScholar authors observe the same about long-tail knowledge in general:
models do worst on what they saw least.

There is a strange detail.
[Agrawal et al.](https://aclanthology.org/2024.findings-eacl.62/) questioned
models about their own references and found that GPT‑4 and the others "often
produce inconsistent author lists for hallucinated references" while they "also
often accurately recall the authors of real references". Their reading is that
the root of the issue may lie "in the generation (i.e., decoding) process"
rather than in what the model knows. The model is not lying to you. It is
completing a pattern, and nothing in that process stops to check.

## Why search alone does not fix it

The obvious fix is to let the model search before it cites. That helps, by
about half.
[What Should I Cite?](https://luneresearch.com/papers/0d10d695-c86b-4812-8a30-5ccf007953d4) (WWW 2026)
measured citations to papers that do not exist with and without retrieval.
Retrieval took GPT‑5 from 23.7% to 13.8%, Claude Sonnet 4 from 25.9% to 10.6%,
and Gemini 2.5 Pro from 22.8% to 11.5%. Better, and still one in ten.

The larger problem is the second failure mode: the paper is real, but it does
not say what your sentence claims. The authors of
[HALoGEN](https://luneresearch.com/papers/f9e6eacf-1ac5-4dea-9a5f-f7266f2aca15) (ACL 2025) warn that
"even if references themselves are not hallucinated, LLMs may still attribute
incorrect claims to them." The measurements agree:

- [Liu et al.](https://aclanthology.org/2023.findings-emnlp.467/) (Findings of
  EMNLP 2023) audited four generative search engines. Only 51.5% of their
  sentences were fully supported by their citations, and 74.5% of citations
  supported the sentence they were attached to.
- [ALCE](https://luneresearch.com/papers/904b10e7-a0c0-460e-9874-85c142af3eb5) (EMNLP 2023) found that
  "on the ELI5 dataset, even the best models lack complete citation support 50%
  of the time."
- [Magesh et al.](https://onlinelibrary.wiley.com/doi/10.1111/jels.12413)
  (Journal of Empirical Legal Studies 2025) tested commercial legal research
  tools built on retrieval. They hallucinated on 17% to 33% of queries, and the
  definition counts any answer that "falsely asserts that a source supports a
  statement".
- [DeepTRACE](https://luneresearch.com/papers/b266c5a6-45e8-4b68-9691-70f2fa27623a) (ICLR 2026) audited
  deep research products as of August 2025 and measured citation accuracy from
  40.3% for Gemini Deep Research to 79.1% for GPT‑5 Deep Research.

Retrieval gets a real paper into the context. It does not make the model read
the paper carefully, and it does not check that the sentence it wrote matches
the passage it cites.

## What actually works

Four design choices address both failure modes. Lune is built on all four, but
none of them is specific to Lune, and you can hold any tool to the same list.

**A closed library.** Every paper Lune returns is a real record from a known
corpus, with a stable identifier and a page you can open, so a reference taken
from the results is one you can trace. No tool can stop a model from adding a
reference of its own, which is why the check is simple: any reference in the
draft that does not match a returned record is the one to cut. Lune's corpus is
papers from top computer science venues ranked by CORE and CCF, and nothing
from the open web.

**Full text, not abstracts.** The OpenScholar team found that even when GPT‑4o
cited real papers, "most of them are not substantiated by the corresponding
abstracts". Abstracts summarize. The numbers, the conditions and the caveats
that decide whether a claim holds are in the body, so a system that only reads
abstracts is guessing at the rest.

**Checking each claim against a quoted passage.** The step that catches the
second failure mode is verification at the level of the sentence: find the
passage that bears on the claim, quote it exactly, and flag the claim when the
evidence contradicts it or cannot be found. OpenScholar's own analysis credits
"reranking, self-feedback and verification" for its results. Lune's
`verify_claims` answers supported, unsupported or insufficient evidence for
each claim, and whenever a verdict comes with a quote, the server has confirmed
that the quote occurs in the retrieved text.

**Permission to say "not found."** A tool that must always answer will
eventually invent. Testing chat assistants on paper search,
[PaperAsk](https://luneresearch.com/papers/cf1c6daa-ac71-4903-a0be-a9b5c80639fe) (WWW 2026) found that
"ChatGPT often withholds responses rather than risk errors, whereas Gemini
produces fluent but fabricated answers." Abstaining is the better failure. Lune
flags a search as low confidence when nothing relevant clears the bar, so the
agent can say the corpus does not cover a question instead of stretching a weak
match.

## A check you can run before you submit

You do not need any particular tool for this, only the discipline to open what
you cite.

1. Resolve every DOI and arXiv ID, and compare the title, authors, venue and
   year with your entry. A real paper with a wrong detail is still an error: in
   one 2025 study, 45% of the real citations GPT‑4o produced contained one, most
   often an incorrect or invalid DOI.
2. Search each title in quotation marks. A title that returns nothing is a red
   flag, especially for a recent paper.
3. For every sentence that carries a citation, find the passage in the cited
   paper that supports it. If you cannot find one, rewrite the sentence or drop
   the citation.
4. Check obscure papers hardest. Models get highly cited papers right more
   often than rarely cited ones.
5. Treat a model's reference list as leads to check, never as a bibliography.

ICLR, NeurIPS and ACL all hold the authors responsible, whoever typed the
reference. The fix is not to stop using AI for literature work. It is to use
tools that can only cite what they can open, and to check the rest yourself.
For a worked example of that workflow end to end, see
[how to run a literature review in Claude Code](https://luneresearch.com/blog/literature-review-with-claude-code).

## Sources

- ACL 2026 program chairs,
  [statement on desk rejecting papers with hallucinated references](https://2026.aclweb.org/acl_statement/)
- ICLR,
  [a retrospective on the ICLR 2026 review process](https://blog.iclr.cc/2026/03/31/a-retrospective-on-the-iclr-2026-review-process/)
- NeurIPS,
  [2026 main track handbook](https://neurips.cc/Conferences/2026/MainTrackHandbook)
- Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023),
  [opinion and order on sanctions](https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf)
- Walters and Wilder, "Fabrication and errors in the bibliographic citations
  generated by ChatGPT,"
  [Scientific Reports 13, 14045 (2023)](https://www.nature.com/articles/s41598-023-41032-5)
- Chelli et al., "Hallucination Rates and Reference Accuracy of ChatGPT and Bard
  for Systematic Reviews,"
  [Journal of Medical Internet Research 26, e53164 (2024)](https://www.jmir.org/2024/1/e53164)
- Agrawal et al., "Do Language Models Know When They're Hallucinating
  References?,"
  [Findings of EACL 2024](https://aclanthology.org/2024.findings-eacl.62/)
- Asai et al., "Synthesizing scientific literature with retrieval-augmented
  language models,"
  [Nature 650, 857 to 863 (2026)](https://www.nature.com/articles/s41586-025-10072-4)
- Tang et al., "Large Language Models for Automated Literature Review,"
  [EMNLP 2025](https://luneresearch.com/papers/e84b5dc7-47b0-4cbe-8f6d-a8da1b37040a)
- Linardon et al., "Influence of Topic Familiarity and Prompt Specificity on
  Citation Fabrication in Mental Health Research Using Large Language Models,"
  [JMIR Mental Health 12, e80371 (2025)](https://mental.jmir.org/2025/1/e80371)
- Zheng et al., "What Should I Cite? A RAG Benchmark for Academic Citation
  Prediction,"
  [WWW 2026](https://luneresearch.com/papers/0d10d695-c86b-4812-8a30-5ccf007953d4)
- Ravichander et al., "HALoGEN: Fantastic LLM Hallucinations and Where to Find
  Them," [ACL 2025](https://luneresearch.com/papers/f9e6eacf-1ac5-4dea-9a5f-f7266f2aca15)
- Liu, Zhang and Liang, "Evaluating Verifiability in Generative Search Engines,"
  [Findings of EMNLP 2023](https://aclanthology.org/2023.findings-emnlp.467/)
- Gao et al., "Enabling Large Language Models to Generate Text with Citations,"
  [EMNLP 2023](https://luneresearch.com/papers/904b10e7-a0c0-460e-9874-85c142af3eb5)
- Magesh et al., "Hallucination-Free? Assessing the Reliability of Leading AI
  Legal Research Tools,"
  [Journal of Empirical Legal Studies 22, 216 to 242 (2025)](https://onlinelibrary.wiley.com/doi/10.1111/jels.12413)
- Venkit et al., "DeepTRACE: Auditing Deep Research AI Systems for Tracking
  Reliability Across Citations and Evidence,"
  [ICLR 2026](https://luneresearch.com/papers/b266c5a6-45e8-4b68-9691-70f2fa27623a)
- Wu et al., "PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper
  Search and Reading,"
  [WWW 2026](https://luneresearch.com/papers/cf1c6daa-ac71-4903-a0be-a9b5c80639fe)
