On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-Collapse
Alkis Kalavasis, Anay Mehrotra, Grigoris Velegkas
Abstract
Specifying all desirable properties of a language model is challenging, but certain requirements seem essential for any good model. Given samples drawn from an unknown language, the trained model should (1) produce valid strings that have not been seen in the training data, and (2) be expressive enough to capture the full richness of the language. Otherwise, if the language model outputs invalid strings, it "hallucinates," and if it fails to capture the full range of the language, it suffers from "mode collapse." In this paper, we ask whether it is possible for a language model to meet both of these requirements. We investigate this question within a statistical setting of language generation, building on the seminal works of Gold [Gol67, Inf. Control], Angluin [Ang79, STOC], and Angluin [Ang88, Tech. Report]. In this setting, the language model is presented with randomly sampled strings from a distribution supported on an unknown language K, which is only known to belong to a possibly infinite collection of candidate languages. The goal of the model is to generate unseen strings from this target language. We say that the language model generates from K with consistency and breadth if, as the size of the training set increases, the set of strings it can output converges to the set of all unseen strings in K. Kleinberg and Mullainathan [KM24, NeurIPS ] posed an open question of whether consistency and breadth in language generation are both possible. We answer this question negatively: for a large class of language models -including next-token-prediction-based models -this is impossible for most collections of candidate languages. This contrasts with the recent positive result of Kleinberg and Mullainathan [KM24, NeurIPS], which demonstrated that consistent generation, without requiring breadth, is possible for any countable collection of candidate languages. Our finding highlights that generation with breadth is fundamentally different from generation without breadth. As a byproduct of our result, we also examine how many samples are required for generation with or without breadth, establishing near-tight bounds on the "learning curves" for generation in the statistical framework of Bousquet, Hanneke, Moran, van Handel, and Yehudayoff [BHM+21, STOC]. Finally, our results also give some hope for consistent generation with breadth: it is achievable for any countable collection of languages when negative examples -in the form of strings outside of K -are available in addition to strings inside of K. This suggests that feedback in post-training, which encodes negative examples, can be crucial in reducing hallucinations while also limiting mode collapse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57b66b37-a3da-4dbb-ab6a-ab070a1062bfCited by top-tier papers17
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han et al.NeurIPS 2025 · 73 citations
- On Union-Closedness of Language GenerationSteve Hanneke, Amin Karbasi, Anay Mehrotra, Grigoris VelegkasNeurIPS 2025 · 17 citations
- Language Generation and Identification from Partial Enumeration: Tight Density Bounds and Topological CharacterizationsJon M. Kleinberg, Fan WeiSTOC 2026 · 14 citations
- Probably Approximately Precision and Recall LearningLee Cohen, Yishay Mansour, Shay Moran, Han ShaoNeurIPS 2025 · 8 citations
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 4 citations
Builds on28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Seeing What a GAN Cannot GenerateDavid Bau, Jun-Yan Zhu, Jonas Wulff, William S. Peebles et al.ICCV 2019 · 342 citations
Related papers
- Language Generation in the LimitJon M. Kleinberg, Sendhil MullainathanNeurIPS 2024 · 45 citations
- Density Measures for Language GenerationJon M. Kleinberg, Fan WeiFOCS 2025 · 1 citation
- Language Generation in the Limit: Complexity Barriers and Implications for LearningMarcelo Arenas, Pablo Barcelo, Luis Cofré, Alexander KozachinskiyICML 2026
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationAdam BouyamournEMNLP 2023 · 9 citations
- Characterizing the Effect of Noise in Language Generation in the LimitAaron Li, Ian ZhangICML 2026 · 4 citations
