Scaling Laws of RoPE-based Extrapolation
Xiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu, Dahua Lin
Abstract
The extrapolation capability of Large Language Models (LLMs) based on Rotary Position Embedding is currently a topic of considerable interest. The mainstream approach to addressing extrapolation with LLMs involves modifying RoPE by replacing 10000, the rotary base of in the original RoPE, with a larger value and providing longer fine-tuning text. In this work, we first observe that fine-tuning a RoPE-based LLM with either a smaller or larger base in pre-training context length could significantly enhance its extrapolation performance. After that, we propose Scaling Laws of RoPE-based Extrapolation, a unified framework from the periodic perspective, to describe the relationship between the extrapolation performance and base value as well as tuning context length. In this process, we also explain the origin of the RoPE-based extrapolation issue by critical dimension for extrapolation. Besides these observations and analyses, we achieve extrapolation up to 1 million context length within only 16K training length on LLaMA2 7B and 13B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ee54aa3-2ed6-4b1e-ae97-ad940ce7b93eCited by top-tier papers57
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentHongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen et al.ICLR 2026 · 231 citations
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
- FiT: Flexible Vision Transformer for Diffusion ModelZeyu Lu, Zidong Wang, Di Huang, Chengyue Wu et al.ICML 2024 · 83 citations
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
Builds on18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
Related papers
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang et al.NeurIPS 2024 · 56 citations
- Frequency Bands in RoPE: Base Frequency and Context Length Shape the Interpolation-Extrapolation Trade-offYui Oka, Itsumi Saito, Kyosuke Nishida, Kuniko SaitoICLR 2026
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Extending Context Window of Large Language Models from a Distributional PerspectiveYingsheng Wu, Yuxuan Gu, Xiaocheng Feng, Weihong Zhong et al.EMNLP 2024 · 1 citation
- CLEX: Continuous Length Extrapolation for Large Language ModelsGuanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang et al.ICLR 2024 · 39 citations
