Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
Hadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu, Cem Koc, Vaishaal Shankar, Oncel Tuzel
摘要
Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences. These datasets are created by randomly concatenating documents of various lengths and then chunking them into sequences of a predetermined target length (concat-and-chunk). Recent attention implementations mask cross-document attention, reducing the effective length of a chunk of tokens. Additionally, training on long sequences becomes computationally prohibitive due to the quadratic cost of attention. In this study, we introduce dataset decomposition, a novel variable sequence length training technique, to tackle these challenges. We decompose a dataset into a union of buckets, each containing sequences of the same size extracted from a unique document. During training, we use variable sequence length and batch-size, sampling simultaneously from all buckets with a curriculum. In contrast to the concat-and-chunk baseline, which incurs a fixed attention cost at every step of training, our proposed method incurs a computational cost proportional to the actual document lengths at each step, resulting in significant savings in training time. We train an 8k context-length 1B model at the same cost as a 2k context-length model trained with the baseline approach. Experiments on a web-scale corpus demonstrate that our approach significantly enhances performance on standard language evaluations and long-context benchmarks, reaching target accuracy with up to 6x faster training compared to the baseline. Our method not only enables efficient pretraining on long sequences but also scales effectively with dataset size. Lastly, we shed light on a critical yet less studied aspect of training large language models: the distribution and curriculum of sequence lengths, which results in a non-negligible difference in performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof 等ICLR 2026 · 被引用 11 次
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught ReasonersReiss Koh, Wonbeen Oh, Jaein Jang, Minhyung Lee 等NeurIPS 2025 · 被引用 8 次
- SkyLadder: Better and Faster Pretraining via Context Window SchedulingTongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen 等NeurIPS 2025 · 被引用 6 次
- Compute-Optimal Quantization-Aware TrainingAleksandr Dremov, David Grangier, Angelos Katharopoulos, Awni HannunICLR 2026 · 被引用 4 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- Efficient Long Context Fine-tuning with Chunk FlowXiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang 等ICML 2025
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language ModelsBoheng Sheng, Jiacheng Yao, Meicong Zhang, Guoxiu HeACL 2025 · 被引用 9 次
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong 等ICML 2024 · 被引用 68 次
- Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language ModelsJunfeng Tian, Da Zheng, Yang Chen, Rui Wang 等ACL 2025 · 被引用 8 次
- Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental ChunkZhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin 等EMNLP 2024
