Base of RoPE Bounds Context Length
Mingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang, Hongyu Lin, Xianpei Han, Weipeng Chen
摘要
Position embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation matrix, has been the de facto choice for position embedding in many LLMs, such as the Llama series. RoPE has been further utilized to extend long context capability, which is roughly based on adjusting the base parameter of RoPE to mitigate out-of-distribution (OOD) problems in position embedding. However, in this paper, we find that LLMs may obtain a superficial long-context ability based on the OOD theory. We revisit the role of RoPE in LLMs and propose a novel property of long-term decay, we derive that the base of RoPE bounds context length: there is an absolute lower bound for the base value to obtain certain context length capability. Our work reveals the relationship between context length and RoPE base both theoretically and empirically, which may shed light on future long context training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Test-Time Training Done RightTianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang 等ICLR 2026 · 被引用 127 次
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin 等NeurIPS 2025 · 被引用 51 次
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMsXiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang 等AAAI 2026 · 被引用 30 次
- Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language ModelsJunfeng Tian, Da Zheng, Yang Chen, Rui Wang 等ACL 2025 · 被引用 8 次
- SkyLadder: Better and Faster Pretraining via Context Window SchedulingTongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
相关 Paper
- Scaling Laws of RoPE-based ExtrapolationXiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu 等ICLR 2024 · 被引用 130 次
- Frequency Bands in RoPE: Base Frequency and Context Length Shape the Interpolation-Extrapolation Trade-offYui Oka, Itsumi Saito, Kyosuke Nishida, Kuniko SaitoICLR 2026
- HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationYuhan Chen, Ang Lv, Jian Luan, Bin Wang 等ACL 2025
- HiRoPE: Length Extrapolation for Code Models Using Hierarchical PositionKechi Zhang, Ge Li, Huangzhao Zhang, Zhi JinACL 2024
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language ModelsGuangxin He, Shen Nie, Fengqi Zhu, Yuankang Zhao 等ICLR 2026 · 被引用 13 次
