Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
Maciej Wolczyk, Bartosz Wójcik, Klaudia Balazy, Igor T. Podolak, Jacek Tabor, Marek Smieja, Tomasz Trzcinski
摘要
The problem of reducing processing time of large deep learning models is a fundamental challenge in many real-world applications. Early exit methods strive towards this goal by attaching additional Internal Classifiers (ICs) to intermediate layers of a neural network. ICs can quickly return predictions for easy examples and, as a result, reduce the average inference time of the whole model. However, if a particular IC does not decide to return an answer early, its predictions are discarded, with its computations effectively being wasted. To solve this issue, we introduce Zero Time Waste (ZTW), a novel approach in which each IC reuses predictions returned by its predecessors by (1) adding direct connections between ICs and (2) combining previous outputs in an ensemble-like manner. We conduct extensive experiments across various datasets and architectures to demonstrate that ZTW achieves a significantly better accuracy vs. inference time trade-off than other recently proposed early exit methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Fast yet Safe: Early-Exiting with Risk ControlMetod Jazbec, Alexander Timans, Tin Hadzi Veljkovic, Kaspar Sakmann 等NeurIPS 2024 · 被引用 35 次
- LGViT: Dynamic Early Exiting for Accelerating Vision TransformerGuanyu Xu, Jiawei Hao, Li Shen, Han Hu 等ACM MM 2023 · 被引用 33 次
- Towards Anytime Classification in Early-Exit Architectures by Enforcing Conditional MonotonicityMetod Jazbec, James Urquhart Allingham, Dan Zhang, Eric T. NalisnickNeurIPS 2023 · 被引用 21 次
- Jointly-Learned Exit and Inference for a Dynamic Neural NetworkFlorence Regol, Joud Chataoui, Mark CoatesICLR 2024 · 被引用 17 次
- Adaptive Computation Modules: Granular Conditional Computation for Efficient InferenceBartosz Wójcik, Alessio Devoto, Karol Pustelnik, Pasquale Minervini 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper5
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song 等NeurIPS 2020 · 被引用 179 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
- Resolution Adaptive Networks for Efficient InferenceLe Yang, Yizeng Han, Xi Chen, Shiji Song 等CVPR 2020
相关 Paper
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu 等AAAI 2023 · 被引用 40 次
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang 等AAAI 2024 · 被引用 26 次
- BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as ExpertsDivya Jyoti Bajpai, Manjesh Kumar HanawalICLR 2025
- You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language ModelShengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang 等CVPR 2023
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 被引用 21 次
