DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search
Lei Yang, Shaoyang Xu, Jianxiang Peng, Shaolin Zhu, Deyi Xiong
Abstract
Large language models (LLMs) based on the Transformer architecture usually have their context length limited due to the high training cost. Recent advancements extend the context window by adjusting the scaling factors of RoPE and fine-tuning. However, suboptimal initialization of these factors results in increased fine-tuning costs and reduced performance at target length. To address these challenges, we propose a novel RoPE-based fine-tuning framework that diverges from conventional scaling factors search. Specifically, we present a Divide-and-Conquer Incremental Search (DCIS) algorithm that strategically determines the better scaling factors. Further finetuning with the identified scaling factors effectively extends the context window of LLMs. Empirical results demonstrate that our methodology not only mitigates performance decay at extended target lengths but also allows the model to fine-tune on short contexts and generalize to long contexts, thereby reducing the cost of fine-tuning. The scaling factors obtained through DCIS can even perform effectively without fine-tuning. Further analysis of the search space reveals that DCIS achieves twice the search efficiency compared to other methods. We also examine the impact of the non-strictly increasing scaling factors utilized in DCIS and evaluate the general capabilities of LLMs across various context lengths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5bac656-5bf1-4aaf-bbfc-91e1252ec2e2Cited by top-tier papers2
- From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for TibetanLei Yang, Leiyu Pan, Bojian Xiong, Renren Jin et al.ACL 2026 · 5 citations
- LongRoPE2: Near-Lossless LLM Context Window ScalingNing Shang, Li Lyna Zhang, Siyuan Wang, Gaokai Zhang et al.ICML 2025
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
Related papers
- PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingDawei Zhu, Nan Yang, Liang Wang, Yifan Song et al.ICLR 2024 · 110 citations
- Scaling Laws of RoPE-based ExtrapolationXiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu et al.ICLR 2024 · 130 citations
- Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationLinda He, Jue Wang, Maurice Weber, Shang Zhu et al.ICLR 2025
- AdaRoPE: Not All Attention Heads Should Rotate and Scale EquallyShaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen et al.ICML 2026
- RoBSA: RoPE-based Blockwise Sparse Multi-head Latent AttentionXinyu Shi, Kairong Luo, Zhen Zheng, Wenguang ChenACL 2026
