USENIX ATC2025顶会
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
Weiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng, Haibin Lin, Minlan Yu
摘要
Multimodal large language models (MLLMs) have extended the success of large language models (LLMs) to multiple data types, such as image, text and audio, achieving significant performance in various domains, including multimodal translation, visual question answering and content generation. Nonetheless, existing systems are inefficient to train MLLMs due to substantial GPU bubbles caused by the heterogeneous modality models and complex data dependencies in 3D parallelism. This paper proposes Optimus, a distributed MLLM training system that reduces end-to-end MLLM training time. Optimus is based on our principled analysis that scheduling the encoder computation within the LLM bubbles can reduce bubbles in MLLM training. To make scheduling encoder computation possible for all GPUs, Optimus searches the separate parallel plans for encoder and LLM, and adopts a bubble scheduling algorithm to enable exploiting LLM bubbles without breaking the original data dependencies in the MLLM model architecture. We further decompose encoder layer computation into a series of kernels, and analyze the common bubble pattern of 3D parallelism to carefully optimize the sub-millisecond bubble scheduling, minimizing the overall training time. Our experiments in a production cluster show that Optimus accelerates MLLM training by 20.5%-21.3% with ViT-22B and GPT-175B model over 3072 GPUs compared to baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline OptimizationHyeonjun An, Sihyun Kim, Chaerim Lim, Hyunjoon Kim 等SIGMOD 2026 · 被引用 1 次
- MegaScale-Data: Scaling DataLoader for Multisource Large Foundation Model TrainingJuntao Zhao, Qi Lu, Wei Jia, Borui Wan 等EuroSys 2026
- OmniScale: Scaling Any Modality Model Training with Model-Centric Distributed Recipe ZooQianli Ma, Yaowei Zheng, Zhelun Shi, Zhongkai Zhao 等AAAI 2026
- MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in ProductionChunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng 等EuroSys 2026
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park 等ICLR 2026
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
相关 Paper
- Accelerating Multi-modal LLM Training with Adaptive Model Placement and ParallelizationYiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen ChuINFOCOM 2026
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu 等NeurIPS 2025 · 被引用 2 次
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 被引用 4 次
- Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUsZhicheng Li, Jiacheng Zhao, Yangyu Zhang, Zhaolin Duan 等ISCA 2026
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsZhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li 等NeurIPS 2025 · 被引用 5 次
