Innovation: An Almost Characterization of Hallucination
Nishant Pratim Das, Piyush Srivastava
Abstract
Hallucination is a central limitation of large language models (LLMs), and substantial effort has been devoted to understanding and mitigating it. Towards this, Kalai and Vempala (STOC 2024) introduced a probabilistic framework formalizing calibration and hallucination, and showed that, with high probability, calibrated LLMs hallucinate roughly at the rate of the "missing mass", a measure of how incomplete the training data is relative to its source. This raises two fundamental questions: (i) what property of a calibrated LLM makes hallucinations unavoidable? and (ii) can hallucinations be avoided by giving up calibration? We answer these questions by introducing a simpler property we call innovation that measures the tendency of a model to produce outputs outside the training data. We show that innovation is implied by the condition for hallucination identified by Kalai and Vempala, and, further, that it is an almost characterization of hallucination: hallucination implies innovation, and conversely, innovation implies hallucination with high probability. We also provide lower bounds on the hallucination rate based on the "innovation rate", and by relating innovation rate back to missing mass, we obtain new hallucination rate lower bounds based on missing mass that extend the results of Kalai and Vempala.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 767200e5-6cd9-4597-bdeb-17602fd66d5fBuilds on7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Calibrated Language Models Must HallucinateAdam Tauman Kalai, Santosh S. VempalaSTOC 2024 · 58 citations
- Language Generation in the LimitJon M. Kleinberg, Sendhil MullainathanNeurIPS 2024 · 45 citations
- On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-CollapseAlkis Kalavasis, Anay Mehrotra, Grigoris VelegkasSTOC 2025 · 2 citations
- Density Measures for Language GenerationJon M. Kleinberg, Fan WeiFOCS 2025 · 1 citation
Related papers
- I Don't Know: Explicit Modeling of Uncertainty with an [IDK] TokenRoi Cohen, Konstantin Dobler, Eden Biran, Gerard de MeloNeurIPS 2024 · 35 citations
- To Believe or Not to Believe Your LLM: Iterative Prompting for Estimating Epistemic UncertaintyYasin Abbasi-Yadkori, Ilja Kuzborskij, András György, Csaba SzepesváriNeurIPS 2024
- HalluLens: LLM Hallucination BenchmarkYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn et al.ACL 2025
- Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human EvaluationsPengqi Li, Lizhong Ding, Zhehao Zhou, Chunhui Zhang et al.ACL 2026
- Enhancing Hallucination Detection through Noise InjectionLitian Liu, Reza Pourreza, Sunny Panchal, Apratim Bhattacharyya et al.ICLR 2026 · 19 citations
