Confabulation: The Surprising Value of Large Language Model Hallucinations
Peiqi Sui, Eamon Duede, Sophie Wu, Richard Jean So
Abstract
This paper presents a systematic defense of large language model (LLM) hallucinations or 'confabulations' as a potential resource instead of a categorically negative pitfall. The standard view is that confabulations are inherently problematic and AI research should eliminate this flaw. In this paper, we argue and empirically demonstrate that measurable semantic characteristics of LLM confabulations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication. In other words, it has potential value. Specifically, we analyze popular hallucination benchmarks and reveal that hallucinated outputs display increased levels of narrativity and semantic coherence relative to veridical outputs. This finding reveals a tension in our usually dismissive understandings of confabulation. It suggests, counter-intuitively, that the tendency for LLMs to confabulate may be intimately associated with a positive capacity for coherent narrative-text generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdf0ee5c-9f52-4afa-9891-2b90e826a19fCited by top-tier papers10
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran et al.VLDB 2025 · 62 citations
- Thinking beyond the anthropomorphic paradigm benefits LLM researchLujain Ibrahim, Myra ChengACL 2026 · 13 citations
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 6 citations
- KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive ReasoningPeiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy et al.ACL 2025 · 6 citations
- Critical Confabulation: Can LLMs Hallucinate for Social Good?Peiqi Sui, Eamon Duede, Hoyt Long, Richard Jean SoICLR 2026 · 2 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- Narrative Theory for Computational Narrative UnderstandingAndrew Piper, Richard Jean So, David BammanEMNLP 2021 · 62 citations
- Calibrated Language Models Must HallucinateAdam Tauman Kalai, Santosh S. VempalaSTOC 2024 · 58 citations
Related papers
- Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language ModelsChaoya Jiang, Hongrui Jia, Mengfan Dong, Wei Ye et al.ACM MM 2024 · 19 citations
- CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language ModelsXiaqiang Tang, Jian Li, Keyu Hu, Nan Du et al.ACL 2025 · 3 citations
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 10 citations
- HILL: A Hallucination Identifier for Large Language ModelsFlorian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble et al.CHI 2024 · 67 citations
- Where Confabulation Lives: Latent Feature Discovery in LLMsThibaud Ardoin, Yi Cai, Gerhard WunderEMNLP 2025 · 1 citation
