An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding
Tong Wu, Yanpeng Zhao, Zilong Zheng
Abstract
Recently, many methods have been developed to extend the context length of pre-trained large language models (LLMs), but they often require fine-tuning at the target length () and struggle to effectively utilize information from the middle part of the context. To address these issues, we propose ontinuity-elativity indxing with gussian iddle (), which interpolates positional encodings by manipulating position indices. Apart from being simple, is training-efficient: it only requires fine-tuning at the pre-trained context window (e.g., Llama 2-4K) and can extend LLMs to a much longer target context length (e.g., 256K). To ensure that the model focuses more on the information in the middle, we introduce a truncated Gaussian to encourage sampling from the middle part of the context during fine-tuning, thus alleviating the"Lost-in-the-Middle"problem faced by long-context LLMs. Experimental results show that successfully extends LLMs to the target length for both Base and Chat versions of with"Never Miss A Beat". Our code is publicly available at https://github.com/bigai-nlco/cream.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef5a584b-3dec-4f81-baf5-472a5d18d019Cited by top-tier papers10
- Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingYoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo CetinICLR 2026 · 13 citations
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement LearningTong Wu, Michael Liu, Jun Bai, Zixia Jia et al.ICML 2026 · 12 citations
- SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference OptimizationHuashan Sun, Shengyi Liao, Yansen Han, Yu Bai et al.ICLR 2026 · 9 citations
- DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question AnsweringJiakai Li, Rongzheng Wang, Yizhuo Ma, Shuang Liang et al.NeurIPS 2025 · 8 citations
- Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language ModelsJinman Zhao, Erxue Min, Hui Wu, Ziheng Li et al.AAAI 2026 · 1 citation
Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
Related papers
- PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingDawei Zhu, Nan Yang, Liang Wang, Yifan Song et al.ICLR 2024 · 110 citations
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
- CLEX: Continuous Length Extrapolation for Large Language ModelsGuanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang et al.ICLR 2024 · 39 citations
- LongRecipe: Recipe for Efficient Long Context Generalization in Large Language ModelsZhiyuan Hu, Yuliang Liu, Jinman Zhao, Suyuchen Wang et al.ACL 2025
- LongRoPE2: Near-Lossless LLM Context Window ScalingNing Shang, Li Lyna Zhang, Siyuan Wang, Gaokai Zhang et al.ICML 2025
