AutoByte: Automatic Configuration for Optimal Communication Scheduling in DNN Training
Yiqing Ma, Hao Wang, Yiming Zhang, Kai Chen
Abstract
ByteScheduler partitions and rearranges tensor transmissions to improve the communication efficiency of distributed Deep Neural Network (DNN) training. The configuration of hyper-parameters (i.e., the partition size and the credit size) is critical to the effectiveness of partitioning and rearrangement. Currently, ByteScheduler adopts Bayesian Optimization (BO) to find the optimal configuration for the hyper-parameters beforehand. In practice, however, various runtime factors (e.g., worker node status and network conditions) change over time, making the statically-determined one-shot configuration result suboptimal for real-world DNN training.
To address this problem, we present a real-time configuration method (called AutoByte) that automatically and timely searches the optimal hyper-parameters as the training systems dynamically change. AutoByte extends the ByteScheduler framework with a meta-network, which takes the system's runtime statistics as its input and outputs predictions for speedups under specific configurations. Evaluation results on various DNN models show that AutoByte can dynamically tune the hyper-parameters with low resource usage, and deliver up to 33.2% higher performance than the best static configuration in ByteScheduler.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingYiding Wang, Decang Sun, Kai Chen, Fan Lai et al.EuroSys 2023 · 43 citations
- Accelerating Privacy-Preserving Machine Learning With GeniBatchXinyang Huang, Junxue Zhang, Xiaodian Cheng, Hong Zhang et al.EuroSys 2024 · 8 citations
Builds on4
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN TrainingWeiyan Wang, Cengguang Zhang, Liu Yang, Kai Chen et al.INFOCOM 2022 · 14 citations
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
Related papers
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the FlyYuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu et al.ICLR 2021 · 26 citations
- xCCLTuner: Treating xCCL as Black-Box and Automatically TuningChenxu Wang, Wentao Fan, Zhehao Lin, Peirui Cao et al.INFOCOM 2026
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin et al.NSDI 2025 · 27 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
