How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies
Piotr Kubaty, Bartosz Wójcik, Bartlomiej Krzepkowski, Monika Michaluk, Tomasz Trzcinski, Jary Pomponi, Kamil Adamczewski
摘要
Early exits enable the network's forward pass to terminate early by attaching trainable internal classifiers to the backbone network. Existing early-exit methods typically adopt either a joint training approach, where the backbone and exit heads are trained simultaneously, or a disjoint approach, where the heads are trained separately. However, the implications of this choice are often overlooked, with studies typically adopting one approach without adequate justification. This choice influences training dynamics and its impact remains largely unexplored. In this paper, we introduce a set of metrics to analyze early-exit training dynamics and guide the choice of training strategy. We demonstrate that conventionally used joint and disjoint regimes yield suboptimal performance. To address these limitations, we propose a mixed training strategy: the backbone is trained first, followed by the training of the entire multiexit network. Through comprehensive evaluations of training strategies across various architectures, datasets, and early-exit methods, we present the strengths and weaknesses of the early exit training strategies. In particular, we show consistent improvements in performance and efficiency using the proposed mixed strategy. How to Train Your Multi-Exit Model? Analyzing the Impact of Training Strategies network first, then freeze its weights and train the parameters of the newly added ICs in the second, separate phase of training (Teerapittayanon et al., 2016; Liao et al., 2021; Zhou et al., 2020) ("disjoint" regime). To the best of our knowledge, no study compares or explores the relationship between these training regimes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
相关 Paper
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 被引用 26 次
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 被引用 11 次
- Dynamic Perceiver for Efficient Visual RecognitionYizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang 等ICCV 2023 · 被引用 45 次
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang 等AAAI 2024 · 被引用 26 次
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang 等AAAI 2025 · 被引用 3 次
