Language Generation in the Limit: Noise, Loss, and Feedback
Yannan Bai, Debmalya Panigrahi, Ian Zhang
摘要
recently proposed a formal framework called language generation in the limit and showed that given a sequence of example strings from an unknown target language drawn from any countable collection, an algorithm can correctly generate unseen strings from the target language within finite time. This notion of language generation was further refined by Li, Raman, and Tewari (2025), who defined progressively stricter categories called non-uniform and uniform generation within generation in the limit. They showed that a finite union of uniformly generatable collections is generatable in the limit, and asked if the same is true for non-uniform generation and generation in the limit.
Our starting point in this paper is to resolve the question of Li, Raman, and Tewari in the negative: we give a uniformly generatable collection and a non-uniformly generatable collection whose union is not generatable in the limit. We then use facets of this construction to further our understanding of several variants of language generation. The first two, language generation with noise and without samples, were introduced by Raman and Raman (2025) and Li, Raman, and Tewari (2025) respectively. We show the equivalence of these models, for both uniform and non-uniform generation. We also provide a complete characterization of non-uniform noisy generation, complementing the corresponding result of Raman and Raman (2025) for uniform noisy generation. The former paper asked if there is any separation between noisy and non-noisy generation in the limit-we show that such a separation exists even with a single noisy string. Finally, we study the framework of generation with feedback, introduced by Charikar and Pabbaraju (2025), where the algorithm is strengthened by allowing it to ask membership queries. We draw a sharp distinction between finite and infinite queries: we show that the former gives no extra power, but the latter is closed under countable union, making it a strictly more powerful model than language generation without feedback.
In summary, the results in this paper resolve the union-closedness of language generation in the limit, and leverage those techniques (and others) to give precise characterizations for natural variants that incorporate noise, loss, and feedback in language generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Language Generation and Identification from Partial Enumeration: Tight Density Bounds and Topological CharacterizationsJon M. Kleinberg, Fan WeiSTOC 2026 · 被引用 14 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
- Characterizing the Effect of Noise in Language Generation in the LimitAaron Li, Ian ZhangICML 2026 · 被引用 4 次
- Language Identification in the Limit with Computational TraceBinghui Peng, Amin Saberi, Grigoris VelegkasICLR 2026
- Language Generation with Feedback: Queries and MistakesSteve Hanneke, Amin Karbasi, Anay Mehrotra, Grigorios VelegkasICML 2026
它引用的顶会 Paper7
- Calibrated Language Models Must HallucinateAdam Tauman Kalai, Santosh S. VempalaSTOC 2024 · 被引用 58 次
- Language Generation in the LimitJon M. Kleinberg, Sendhil MullainathanNeurIPS 2024 · 被引用 45 次
- On Union-Closedness of Language GenerationSteve Hanneke, Amin Karbasi, Anay Mehrotra, Grigoris VelegkasNeurIPS 2025 · 被引用 17 次
- On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-CollapseAlkis Kalavasis, Anay Mehrotra, Grigoris VelegkasSTOC 2025 · 被引用 2 次
- Density Measures for Language GenerationJon M. Kleinberg, Fan WeiFOCS 2025 · 被引用 1 次
相关 Paper
- Representative Language GenerationCharlotte Peale, Vinod Raman, Omer ReingoldICML 2025
- Language Generation in the Limit: Complexity Barriers and Implications for LearningMarcelo Arenas, Pablo Barcelo, Luis Cofré, Alexander KozachinskiyICML 2026
- Generation from Noisy ExamplesAnanth Raman, Vinod RamanICML 2025
- Balancing Understanding and Generation in Discrete Diffusion ModelsYue Liu, Yuzhong Zhao, Zheyong Xie, Qixiang Ye 等ICML 2026
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
