USENIX ATC2025顶会
FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length Inputs
Hairui Zhao, Qi Tian, Hongliang Li, Zizhong Chen
摘要
Transformer achieves promising results among various deep learning architectures. Training transformer-based models (transformers) typically involves various parallelisms, such as data parallelism and pipeline parallelism (PP). Variable-length datasets have been adopted to facilitate multi-task training of transformers, which degrades training efficiency. Though many efforts have significantly improved the variable-length training, these efforts primarily focus on optimizations within a single iteration. However, substantial fluctuations of computation and memory requirements across iterations can also lead to inefficiency overall due to the static partitioning of distributed frameworks. Thus, this paper proposes FlexPipe from the perspective of a distributed system to enable high throughput variable-length training of transformers. To our knowledge, FlexPipe is the first flexible pipeline framework that dynamically adjusts PP by a live flexibility mechanism without training loss. We introduce a novel problem which aims at maximizing training throughput by adjusting the parallel configurations, along with an efficient heuristic algorithm to solve the problem. Extensive experiments show that Flex-Pipe achieves an average 1.25× training throughput compared to state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal ParallelismZedong Liu, Shenggan Cheng, Guangming Tan, Yang You 等NeurIPS 2025 · 被引用 12 次
- DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved PipelineZhenliang Xue, Hanpeng Hu, Xing Chen, Yimin Jiang 等ASPLOS 2026 · 被引用 1 次
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun 等PPoPP 2026
它引用的顶会 Paper25
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
相关 Paper
- MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale ModelsWeigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang 等DAC 2023 · 被引用 9 次
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu 等PPoPP 2026 · 被引用 3 次
- DynaPipe: Optimizing Multi-task Training through Dynamic PipelinesChenyu Jiang, Zhen Jia, Shuai Zheng, Yida Wang 等EuroSys 2024 · 被引用 10 次
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang 等PPoPP 2025 · 被引用 5 次
- AdaptPipe: Mitigating Runtime Bubbles via Granularity-Adaptive Scheduling under Memory ConstraintsYumeng Cui, Jessie Hui Wang, Najila Liu, Ling Deng 等INFOCOM 2026
