Base of RoPE Bounds Context Length
Mingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang, Hongyu Lin, Xianpei Han, Weipeng Chen
Abstract
Position embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation matrix, has been the de facto choice for position embedding in many LLMs, such as the Llama series. RoPE has been further utilized to extend long context capability, which is roughly based on adjusting the base parameter of RoPE to mitigate out-of-distribution (OOD) problems in position embedding. However, in this paper, we find that LLMs may obtain a superficial long-context ability based on the OOD theory. We revisit the role of RoPE in LLMs and propose a novel property of long-term decay, we derive that the base of RoPE bounds context length: there is an absolute lower bound for the base value to obtain certain context length capability. Our work reveals the relationship between context length and RoPE base both theoretically and empirically, which may shed light on future long context training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 728923b8-55bf-4804-9e3f-957859a31041Cited by top-tier papers21
- Test-Time Training Done RightTianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang et al.ICLR 2026 · 127 citations
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin et al.NeurIPS 2025 · 51 citations
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMsXiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang et al.AAAI 2026 · 30 citations
- Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language ModelsJunfeng Tian, Da Zheng, Yang Chen, Rui Wang et al.ACL 2025 · 8 citations
- SkyLadder: Better and Faster Pretraining via Context Window SchedulingTongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen et al.NeurIPS 2025 · 6 citations
Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
Related papers
- Scaling Laws of RoPE-based ExtrapolationXiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu et al.ICLR 2024 · 130 citations
- Frequency Bands in RoPE: Base Frequency and Context Length Shape the Interpolation-Extrapolation Trade-offYui Oka, Itsumi Saito, Kyosuke Nishida, Kuniko SaitoICLR 2026
- HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationYuhan Chen, Ang Lv, Jian Luan, Bin Wang et al.ACL 2025
- HiRoPE: Length Extrapolation for Code Models Using Hierarchical PositionKechi Zhang, Ge Li, Huangzhao Zhang, Zhi JinACL 2024
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language ModelsGuangxin He, Shen Nie, Fengqi Zhu, Yuankang Zhao et al.ICLR 2026 · 13 citations
