How to Train Long-Context Language Models (Effectively)
Tianyu Gao, Alexander Wettig, Howard Yen, Danqi Chen
Abstract
We study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information. We first establish a reliable evaluation protocol to guide model developmentinstead of perplexity or simple needle-in-a-haystack (NIAH) tests, we use a broad set of long-context downstream tasks, and we evaluate models after SFT as this better reveals long-context abilities. Supported by our robust evaluations, we run thorough experiments to decide the data mix for continued pre-training, the instruction tuning dataset, and many other design choices such as position extrapolation. We find that (1) code repositories and books are excellent sources of long data, but it is crucial to combine them with high-quality short-context data; (2) training with a sequence length beyond the evaluation length boosts long-context performance; (3) for SFT, using only short instruction datasets yields strong performance on long-context tasks. Our final model, ProLong-8B, which is initialized from Llama-3 and trained on 40B tokens, demonstrates state-of-the-art longcontext performance among similarly sized models at a length of 128K. ProLong outperforms Llama-3.1-8B-Instruct on the majority of long-context tasks despite using only 5% as many tokens during long-context training. Additionally, ProLong can effectively process up to 512K tokens, one of the longest context windows of publicly available LMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdf9a9c8-22d0-4cf3-97a6-6aff3f0d066aCited by top-tier papers45
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen et al.ACL 2026 · 57 citations
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin et al.NeurIPS 2025 · 51 citations
- Seq vs Seq: An Open Suite of Paired Encoders and DecodersOrion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin et al.ICLR 2026 · 50 citations
- mmBERT: A Modern Multilingual Encoder with Annealed Language LearningMarc Marone, Orion Weller, William Fleshman, Eugene Yang et al.ICML 2026 · 49 citations
Builds on40
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue et al.ICML 2024 · 204 citations
- When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language ModelsYingming Zheng, Hanqi Li, Kai Yu, Lu ChenEMNLP 2025
- LongRoPE2: Near-Lossless LLM Context Window ScalingNing Shang, Li Lyna Zhang, Siyuan Wang, Gaokai Zhang et al.ICML 2025
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
- Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationLinda He, Jue Wang, Maurice Weber, Shang Zhu et al.ICLR 2025
