Intrinsic Entropy of Context Length Scaling in LLMs
Jingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao, Jenq-Neng Hwang, Lei Li
Abstract
Long Context Language Models have drawn great attention in the past few years. There has been work discussing the impact of long context on Language Model performance: some find that long irrelevant context could harm performance, while some experimentally summarize loss reduction by relevant long context as Scaling Laws. This calls for a more thorough understanding of how long context impacts Language Modeling. In this work, we (1) propose to use 'Intrinsic Entropy' for explaining the impact of context length on language modeling; and (2) conduct experiments on natural language and synthetic data, validating our proposed theoretical assumptions and deductions. Our theoretical framework can provide practical insights such as establishing that training dataset size dictates an optimal context length and bounds context length scaling for certain cases. We hope our work may inspire new long context Language Models, as well as future work studying the physics of Language Models. 1 † CPHOS is an academic non-profit organization. 1 Code for experiments is available at: https://github.com/JingzheShi/ NLPCtlScalingAndBounds . 2 We discuss more about previous work in Appendix J.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou et al.CVPR 2026 · 12 citations
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic VerificationWenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li et al.KDD 2026
- Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy OptimizationYunhan Bu, quan zhang, Zhang Huaping, Guotong Geng et al.ICML 2026
Builds on10
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff et al.NeurIPS 2024 · 135 citations
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 77 citations
- Scaling Law for Time Series ForecastingJingzhe Shi, Qinwei Ma, Huan Ma, Lei LiNeurIPS 2024 · 39 citations
- Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional DataAlexander Havrilla, Wenjing LiaoNeurIPS 2024 · 36 citations
Related papers
- LM: Mutual Information Scaling Law for Long-Context Language ModelingZhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo et al.NeurIPS 2025 · 11 citations
- Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?Jonathan Roberts, Kai Han, Samuel AlbanieICLR 2025
- What is Wrong with Perplexity for Long-context Language Modeling?Lizhe Fang, Yifei Wang, Zhaoyang Liu, Chenheng Zhang et al.ICLR 2025 · 2 citations
- LongCodeU: Benchmarking Long-Context Language Models on Long Code UnderstandingJia Li, Xuyuan Guo, Lei Li, Kechi Zhang et al.ACL 2025
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
