Intrinsic Entropy of Context Length Scaling in LLMs
Jingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao, Jenq-Neng Hwang, Lei Li
摘要
Long Context Language Models have drawn great attention in the past few years. There has been work discussing the impact of long context on Language Model performance: some find that long irrelevant context could harm performance, while some experimentally summarize loss reduction by relevant long context as Scaling Laws. This calls for a more thorough understanding of how long context impacts Language Modeling. In this work, we (1) propose to use 'Intrinsic Entropy' for explaining the impact of context length on language modeling; and (2) conduct experiments on natural language and synthetic data, validating our proposed theoretical assumptions and deductions. Our theoretical framework can provide practical insights such as establishing that training dataset size dictates an optimal context length and bounds context length scaling for certain cases. We hope our work may inspire new long context Language Models, as well as future work studying the physics of Language Models. 1 † CPHOS is an academic non-profit organization. 1 Code for experiments is available at: https://github.com/JingzheShi/ NLPCtlScalingAndBounds . 2 We discuss more about previous work in Appendix J.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic VerificationWenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li 等KDD 2026
- Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy OptimizationYunhan Bu, quan zhang, Zhang Huaping, Guotong Geng 等ICML 2026
它引用的顶会 Paper10
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff 等NeurIPS 2024 · 被引用 135 次
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 被引用 77 次
- Scaling Law for Time Series ForecastingJingzhe Shi, Qinwei Ma, Huan Ma, Lei LiNeurIPS 2024 · 被引用 39 次
- Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional DataAlexander Havrilla, Wenjing LiaoNeurIPS 2024 · 被引用 36 次
相关 Paper
- LM: Mutual Information Scaling Law for Long-Context Language ModelingZhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo 等NeurIPS 2025 · 被引用 11 次
- Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?Jonathan Roberts, Kai Han, Samuel AlbanieICLR 2025
- What is Wrong with Perplexity for Long-context Language Modeling?Lizhe Fang, Yifei Wang, Zhaoyang Liu, Chenheng Zhang 等ICLR 2025 · 被引用 2 次
- LongCodeU: Benchmarking Long-Context Language Models on Long Code UnderstandingJia Li, Xuyuan Guo, Lei Li, Kechi Zhang 等ACL 2025
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
