Density Measures for Language Generation
Jon M. Kleinberg, Fan Wei
摘要
The recent successes of large language models (LLMs) have led to a surge of theoretical research into the properties of language generation. A recent line of work has proposed an abstract view of the question -called language generation in the limit -in which we view language generation as a game played between an adversary and an algorithm: the adversary generates strings from an unknown language K, known only to come from a countable collection of candidate languages, and after observing a finite set of these strings, the algorithm must generate new strings from the language K that it hasn't seen before. This formalism highlights an important tension: the trade-off between validity (that the algorithm should only produce strings that come from the language) and breadth (that the algorithm should be able to produce "many" strings from the language). This validity-breadth trade-off is a central issue in applied work on language generation as well, where it arises in the balance between hallucination, when models generate invalid utterances, and mode collapse, when models only generate from a very restricted set of feasible outputs. Despite its importance, this trade-off has been challenging to study quantitatively.
In this work we develop ways of quantifying this trade-off, by formalizing the notion of breadth through measures of density. Roughly speaking, the density of one language L in another language L ′ is the limiting fraction of strings from L among the strings of L ′ , where we take the limit over longer and longer finite prefixes of L ′ . Existing algorithms for language generation in the limit produce output sets that can have zero density in the true language K, in this asymptotic sense, and this represents an important failure of breadth that might seem necessary in any solution to the problem. We show here that such a failure is not in fact necessary: we provide an algorithm for language generation in the limit whose outputs have strictly positive density in the true language K. We also study the internal representations built by algorithms for this problem -the sequence of hypothesized candidate languages they iterate through as they perform generation -showing a precise sense in which the strongest form of breadth achievable is one that may need to "oscillate" indefinitely between hypothesized representations of high density and low density. Our analysis introduces a novel topology on language families, with notions of convergence and limit points in this topology playing a key role in the analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- On Union-Closedness of Language GenerationSteve Hanneke, Amin Karbasi, Anay Mehrotra, Grigoris VelegkasNeurIPS 2025 · 被引用 17 次
- Language Generation and Identification from Partial Enumeration: Tight Density Bounds and Topological CharacterizationsJon M. Kleinberg, Fan WeiSTOC 2026 · 被引用 14 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
- Characterizing the Effect of Noise in Language Generation in the LimitAaron Li, Ian ZhangICML 2026 · 被引用 4 次
- On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-CollapseAlkis Kalavasis, Anay Mehrotra, Grigoris VelegkasSTOC 2025 · 被引用 2 次
它引用的顶会 Paper6
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 被引用 183 次
- Representational Strengths and Limitations of TransformersClayton Sanford, Daniel J. Hsu, Matus TelgarskyNeurIPS 2023 · 被引用 162 次
- Calibrated Language Models Must HallucinateAdam Tauman Kalai, Santosh S. VempalaSTOC 2024 · 被引用 58 次
- Language Generation in the LimitJon M. Kleinberg, Sendhil MullainathanNeurIPS 2024 · 被引用 45 次
- ALPINE: Unveiling The Planning Capability of Autoregressive Learning in Language ModelsSiwei Wang, Yifei Shen, Shi Feng, Haoran Sun 等NeurIPS 2024 · 被引用 18 次
相关 Paper
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 被引用 36 次
- Innovation: An Almost Characterization of HallucinationNishant Pratim Das, Piyush SrivastavaICML 2026
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 被引用 197 次
- Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models DecodingLifu Tu, Semih Yavuz, Jin Qu, Jiacheng Xu 等EMNLP 2024 · 被引用 2 次
