LM: Mutual Information Scaling Law for Long-Context Language Modeling
Zhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo, Marin Soljacic
Abstract
We present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that we rigorously verify in natural language. We demonstrate that bipartite mutual information captures multi-token interactions distinct from and scaling independently of conventional two-point mutual information, and show that this provides a more complete characterization of the dependencies needed for accurately modeling long sequences. Leveraging this scaling law, we formulate the Long-context Language Modeling (LM) condition, which lower bounds the necessary scaling of a model's history state -- the latent variables responsible for storing past information -- for effective long-context modeling. We validate the framework and its predictions on transformer and state-space models. Our work provides a principled foundation to understand long-context modeling and to design more efficient architectures with stronger long-context capabilities, with potential applications beyond natural language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ad37b0b-8ec6-4ee3-b907-6654c6150d94Cited by top-tier papers5
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM ReasoningChen Qian, Dongrui Liu, Haochen Wen, Zhen Bai et al.NeurIPS 2025 · 63 citations
- Intrinsic Entropy of Context Length Scaling in LLMsJingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao et al.ICLR 2026 · 17 citations
- Correlation Dimension of Autoregressive Large Language ModelsXin Du, Kumiko Tanaka-IshiiNeurIPS 2025 · 2 citations
- Trading Complexity for Expressivity Through Structured Generalized Linear Token MixingErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICML 2026 · 1 citation
- L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information ScalingZhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo et al.ICML 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
Related papers
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Calibration, Entropy Rates, and Memory in Language ModelsMark Braverman, Xinyi Chen, Sham M. Kakade, Karthik Narasimhan et al.ICML 2020 · 48 citations
- Repeat After Me: Transformers are Better than State Space Models at CopyingSamy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran MalachICML 2024 · 176 citations
- Context-level Language Modeling by Learning Predictive Context Embeddingsbeiya dai, Yuliang Liu, Yunchong Song, Daozheng Xue et al.ICML 2026 · 5 citations
- Do Long-Range Language Models Actually Use Long-Range Context?Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit IyyerEMNLP 2021 · 35 citations
