Analysing The Impact of Sequence Composition on Language Model Pre-Training
Yu Zhao, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski, Wei Liu, Piotr Milos, Yuxiang Wu, Pasquale Minervini
摘要
Most language model pre-training frameworks concatenate multiple documents into fixedlength sequences and use causal masking to compute the likelihood of each token given its context; this strategy is widely adopted due to its simplicity and efficiency. However, to this day, the influence of the pre-training sequence composition strategy on the generalisation properties of the model remains underexplored. In this work, we find that applying causal masking can lead to the inclusion of distracting information from previous documents during pre-training, which negatively impacts the performance of the models on language modelling and downstream tasks. In intra-document causal masking, the likelihood of each token is only conditioned on the previous tokens in the same document, eliminating potential distracting information from previous documents and significantly improving performance. Furthermore, we find that concatenating related documents can reduce some potential distractions during pre-training, and our proposed efficient retrieval-based sequence construction method, BM25Chunk, can improve incontext learning (+11.6%), knowledge memorisation (+9.8%), and context utilisation (+7.2%) abilities of language models without sacrificing efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Structured Packing in LLM Training Improves Long Context UtilizationKonrad Staniszewski, Szymon Tworkowski, Sebastian Jaszczur, Yu Zhao 等AAAI 2025 · 被引用 17 次
- Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized WorldsFan Wang, Pengtao Shao, Yiming Zhang, Bo Yu 等NeurIPS 2025 · 被引用 6 次
- SkyLadder: Better and Faster Pretraining via Context Window SchedulingTongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen 等NeurIPS 2025 · 被引用 6 次
- Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan 等ACL 2025 · 被引用 4 次
- A Simple and Effective L_2 Norm-Based Strategy for KV Cache CompressionAlessio Devoto, Yu Zhao, Simone Scardapane, Pasquale MinerviniEMNLP 2024 · 被引用 3 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
相关 Paper
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
- Enhancing Elusive Clues in Knowledge Learning by Contrasting Attention of Language ModelsJian Gao, Xiao Zhang, Miao Li, Ji WuAAAI 2025 · 被引用 1 次
- Learned Meta-Tokens for Language ModelingAlok N. Shah, Khush Gupta, Keshav Ramji, Pratik ChaudhariICLR 2026 · 被引用 2 次
- Dataset Decomposition: Faster LLM Training with Variable Sequence Length CurriculumHadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu 等NeurIPS 2024 · 被引用 40 次
- Fewer Truncations Improve Language ModelingHantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar 等ICML 2024 · 被引用 28 次
