Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search
Zhiqi Lin, Youshan Miao, Guanbin Xu, Cheng Li, Olli Saarikivi, Saeed Maleki, Fan Yang
Abstract
Increasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require carefully planned schedules for performance. However, existing practices often rely on predefined schedules that may not fully exploit the benefits of emerging diverse model-aware operator placement strategies. Handcrafting high-efficiency schedules can be challenging due to the large and varying schedule space. This paper presents Tessel, an automated system that searches for efficient schedules for distributed DNN training and inference for diverse operator placement strategies. To reduce search costs, Tessel leverages the insight that the most efficient schedules often exhibit repetitive pattern (repetend) across different data inputs. This leads to a two-phase approach: repetend construction and schedule completion. By exploring schedules for various operator placement strategies, Tessel significantly improves both training and inference performance. Experiments with representative DNN models demonstrate that Tessel achieves up to 5.5 x training performance speedup and up to 38 % inference latency reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c467e99-3899-444b-9a6e-2c99fa7b77b9Cited by top-tier papers4
- nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning TrainingZhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang et al.OSDI 2024 · 38 citations
- Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismWeijian Liu, Mingzhen Li, Guangming Tan, Weile JiaPPoPP 2025 · 4 citations
- Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignChunyu Xue, Weihao Cui, Quan Chen, Chen Chen et al.EuroSys 2026
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai et al.ICLR 2022 · 950 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
- CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsHanpeng Hu, Junwei Su, Juntao Zhao, Yanghua Peng et al.EuroSys 2024 · 7 citations
- ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsYang Bai, Wenqian Zhao, Shuo Yin, Zixiao Wang et al.EMNLP 2023 · 2 citations
- Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN PruningYue Guan, Changming Yu, Yangjie Zhou, Jingwen Leng et al.ASPLOS 2024 · 6 citations
- Crane: Inter-Layer Scheduling Framework for DNN Inference and Training Co-Support on Tiled ArchitectureYu Gong, Lingyi Huang, Haodong Chang, Rongjian Liang et al.MICRO 2025 · 1 citation
