Structured Packing in LLM Training Improves Long Context Utilization
Konrad Staniszewski, Szymon Tworkowski, Sebastian Jaszczur, Yu Zhao, Henryk Michalewski, Lukasz Kucinski, Piotr Milos
Abstract
Recent advancements in long-context language modeling have attracted significant attention, yet their practical applications often suffer from suboptimal context utilization. To efficiently address this issue, we introduce the Structured Packing for Long Context, SPLiCe, a method that uses retrieval to collate mutually relevant documents into long training samples. We demonstrate that SPLiCe improves performance on long-context tasks, particularly by achieving perfect accuracy on the synthetic Needle in the Haystack benchmark, and effectively mitigating the ‘lost-in-the-middle’ phenomenon often observed in large language models. Notably, these long-context capabilities also extend to realistic downstream tasks, such as Qasper, across multiple model sizes—3B, 7B, and 13B—and are achieved with only brief fine-tuning on 2-6 billion tokens. We supplement these results with a detailed analysis of SPLiCe, examining the impact of hyperparameter choices, the different mixtures and proportions of SPLiCe-generated training data, and the choice of the retriever. We also study the transfer of long-context utilization skills between the modalities. An intriguing finding from our analysis is that training on a corpus of code can enhance performance on natural language tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang et al.NeurIPS 2024 · 42 citations
- DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong et al.ACL 2025 · 12 citations
- Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language ModelsLongze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng et al.ACL 2024 · 4 citations
- Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan et al.ACL 2025 · 4 citations
- Attribute or Abstain: Large Language Models as Long Document AssistantsJan Buchmann, Xiao Liu, Iryna GurevychEMNLP 2024
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
Related papers
- Training with "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-Context TasksYijiong Yu, Yongfeng Huang, Zhixiao Qi, Zhe ZhouAAAI 2025 · 5 citations
- Fewer Truncations Improve Language ModelingHantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar et al.ICML 2024 · 28 citations
- How to Train Long-Context Language Models (Effectively)Tianyu Gao, Alexander Wettig, Howard Yen, Danqi ChenACL 2025
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou et al.ICLR 2024 · 87 citations
- Understanding Synthetic Context Extension via Retrieval HeadsXinyu Zhao, Fangcong Yin, Greg DurrettICML 2025
