A Measure-Theoretic Characterization of Tight Language Models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, Ryan Cotterell
Abstract
Language modeling, a central task in natural language processing, involves estimating a probability distribution over strings. In most cases, the estimated distribution sums to 1 over all finite strings. However, in some pathological cases, probability mass can "leak" onto the set of infinite sequences. In order to characterize the notion of leakage more precisely, this paper offers a measure-theoretic treatment of language modeling. We prove that many popular language model families are in fact tight, meaning that they will not leak in this sense. We also generalize characterizations of tightness proposed in previous works.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37276c68-eb8e-4c64-a7ff-84a7a84f7a3cCited by top-tier papers18
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
- On Affine Homotopy between Language EncodersRobin Chan, Reda Boumasmoud, Anej Svete, Yuxin Ren et al.NeurIPS 2024 · 7 citations
- Principled Gradient-Based MCMC for Conditional Sampling of TextLi Du, Afra Amini, Lucas Torroba Hennigen, Xinyan Velocity Yu et al.ICML 2024 · 6 citations
- MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-EntropiesShiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu et al.ACL 2023 · 5 citations
- Structured Voronoi SamplingAfra Amini, Li Du, Ryan CotterellNeurIPS 2023 · 5 citations
Builds on4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang et al.EMNLP 2020 · 37 citations
- On the Uncomputability of Partition Functions in Energy-Based Sequence ModelsChu-Cheng Lin, Arya D. McCarthyICLR 2022 · 4 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
Related papers
- Are Language Models Any Good at Density Modeling?Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam ChattopadhyayAAAI 2026
- Evaluating Distributional Distortion in Neural Language ModelingBenjamin LeBrun, Alessandro Sordoni, Timothy J. O'DonnellICLR 2022 · 26 citations
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara et al.ICML 2025
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda et al.ACL 2024
