AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline Parallelism
Yan Wang, Chang Si, Kaiming Yang, Zhipeng Zhang, Weijian Liu, Man Yuan, Mingzhen Li, Yong Li, Weile Jia
Abstract
Multi-token prediction (MTP) architecture is widely adopted in LLMs. MTP blocks can be appended to the tail of model to predict additional tokens. However, when training with pipeline parallel, MTP leads to more pipeline bubbles and deteriorates the pipeline efficiency. Based on in-depth analysis of MTP architectures and loss functions, we have identified the parallel nature of the MTP blocks, and leverage it for superior pipeline scheduling. We propose AdaHC, an adaptive pipeline scheduling framework for accelerating LLMs training with MTP block(s). AdaHC splits the output heads into chunks and reassembles the chunks to generate balanced pipeline stages, and performs adaptive activation forwarding to preserve the numerical equivalence. Experimental results show that AdaHC improves the training throughput of SOTA LLMs with diverse MTP configurations by 1.35 on average. This work paves a new direction for practical pipeline training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87f52a02-60b1-487a-b0b2-4068363d549aBuilds on9
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 124 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
Related papers
- AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory ConstraintsYumeng Cui, Jessie Hui Wang, Najila Liu, Ling Deng et al.INFOCOM 2026
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- Pre-Training Curriculum for Multi-Token Prediction in Language ModelsAnsar Aynetdinov, Alan AkbikACL 2025
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu et al.NeurIPS 2025 · 2 citations
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo et al.NeurIPS 2023 · 159 citations
