Empowering Adaptive Early-Exit Inference with Latency Awareness
Xinrui Tan, Hongjia Li, Liming Wang, Xueqing Huang, Zhen Xu
摘要
With the capability of trading accuracy for latency on-the-fly, the technique of adaptive early-exit inference has emerged as a promising line of research to accelerate the deep learning inference. However, studies in this line of research commonly use a group of thresholds to control the accuracy-latency trade-off, where a thorough and general methodology on how to determine these thresholds has not been conducted yet, especially with regard to the common requirements of average inference latency. To address this issue and enable latency-aware adaptive early-exit inference, in the present paper, we approximately formulate the threshold determination problem of finding the accuracy-maximum threshold setting that meets a given average latency requirement, and then propose a threshold determination method to tackle our formulated non-convex problem. Theoretically, we prove that, for certain parameter settings, our method finds an approximate stationary point of the formulated problem. Empirically, on top of various models across multiple datasets (CIFAR-10, CIFAR-100, ImageNet and two time-series datasets), we show that our method can well handle the average latency requirements, and consistently finds good threshold settings in negligible time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Spike-inspired rank coding for fast and accurate recurrent neural networksAlan Jeffares, Qinghai Guo, Pontus Stenetorp, Timoleon MoraitisICLR 2022 · 被引用 17 次
- Window-Based Early-Exit Cascades for Uncertainty Estimation: When Deep Ensembles are More Efficient than Single ModelsGuoxuan Xia, Christos-Savvas BouganisICCV 2023 · 被引用 17 次
它引用的顶会 Paper3
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive InferenceTing-Kuei Hu, Tianlong Chen, Haotao Wang, Zhangyang WangICLR 2020 · 被引用 89 次
相关 Paper
- Rethinking Calibration for Early-Exit Neural NetworksPiotr Kubaty, Filip Szatkowski, Grzegorz Choczyński, Eric Nalisnick 等ICML 2026
- Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationAnand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan 等SOSP 2024 · 被引用 1 次
- Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and ReliabilityDivya Jyoti Bajpai, Manjesh Kumar HanawalNeurIPS 2025 · 被引用 4 次
- Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML ServingYinwei Dai, Rui Pan, Anand P. Iyer, Kai Li 等SOSP 2024 · 被引用 4 次
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu 等AAAI 2023 · 被引用 40 次
