How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
Piotr Kubaty, Bartosz Wójcik, Bartlomiej Krzepkowski, Monika Michaluk, Tomasz Trzcinski, Jary Pomponi, Kamil Adamczewski
Abstract
Early exits enable the network's forward pass to terminate early by attaching trainable internal classifiers to the backbone network. Existing early-exit methods typically adopt either a joint training approach, where the backbone and exit heads are trained simultaneously, or a disjoint approach, where the heads are trained separately. However, the implications of this choice are often overlooked, with studies typically adopting one approach without adequate justification. This choice influences training dynamics and its impact remains largely unexplored. In this paper, we introduce a set of metrics to analyze early-exit training dynamics and guide the choice of training strategy. We demonstrate that conventionally used joint and disjoint regimes yield suboptimal performance. To address these limitations, we propose a mixed training strategy: the backbone is trained first, followed by the training of the entire multiexit network. Through comprehensive evaluations of training strategies across various architectures, datasets, and early-exit methods, we present the strengths and weaknesses of the early exit training strategies. In particular, we show consistent improvements in performance and efficiency using the proposed mixed strategy. How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies network first, then freeze its weights and train the parameters of the newly added ICs in the second, separate phase of training (Teerapittayanon et al., 2016; Liao et al., 2021; Zhou et al., 2020) ("disjoint" regime). To the best of our knowledge, no study compares or explores the relationship between these training regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db6f539e-40a1-477c-9239-ade632c2bbafBuilds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 205 citations
Related papers
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 26 citations
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 11 citations
- Dynamic Perceiver for Efficient Visual RecognitionYizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang et al.ICCV 2023 · 45 citations
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang et al.AAAI 2024 · 26 citations
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang et al.AAAI 2025 · 3 citations
