LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
Zhiyuan Hu, Yuliang Liu, Jinman Zhao, Suyuchen Wang, Yan Wang, Wei Shen, Qing Gu, Anh Tuan Luu, See-Kiong Ng, Zhiwei Jiang, Bryan Hooi
Abstract
Large language models (LLMs) face significant challenges in handling long-context tasks because of their limited effective context window size during pretraining, which restricts their ability to generalize over extended sequences. Meanwhile, extending the context window in LLMs through post-pretraining is highly resource-intensive. To address this, we introduce LongRecipe, an efficient training strategy for extending the context window of LLMs, including impactful token analysis, position index transformation, and training optimization strategies. It simulates long-sequence inputs while maintaining training efficiency and significantly improves the model's understanding of long-range dependencies. Experiments on three types of LLMs show that Lon-gRecipe can utilize long sequences while requiring only 30% of the target context window size, and reduces computational training resource over 85% compared to full sequence training. Furthermore, LongRecipe also preserves the original LLM's capabilities in general tasks. Ultimately, we can extend effective context window of open-source LLMs from 8k to 128k, achieving performance close to GPT-4 with just one day of dedicated training using a single GPU with 80G memory. Our code is released at https://github.com/ zhiyuanhubj/LongRecipe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- QFFT, Question-Free Fine-Tuning for Adaptive ReasoningWanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin et al.NeurIPS 2025 · 26 citations
- UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsXueyan Zhang, Jinman Zhao, Zhifei Yang, Yibo Zhong et al.ACL 2025 · 15 citations
- Safety Alignment via Constrained Knowledge UnlearningZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin et al.ACL 2025 · 8 citations
- FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual LearningYujie Feng, Hao Wang, Jian Li, Xu Chu et al.ACL 2026 · 3 citations
- What is Wrong with Perplexity for Long-context Language Modeling?Lizhe Fang, Yifei Wang, Zhaoyang Liu, Chenheng Zhang et al.ICLR 2025 · 2 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
Related papers
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue et al.ICML 2024 · 204 citations
- LOGO - Long cOntext aliGnment via efficient preference OptimizationZecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu et al.ICML 2025
- PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingDawei Zhu, Nan Yang, Liang Wang, Yifan Song et al.ICLR 2024 · 110 citations
- An Efficient Recipe for Long Context Extension via Middle-Focused Positional EncodingTong Wu, Yanpeng Zhao, Zilong ZhengNeurIPS 2024 · 17 citations
- LLM Maybe LongLM: SelfExtend LLM Context Window Without TuningHongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang et al.ICML 2024 · 167 citations
