Anytime Inference with Distilled Hierarchical Neural Ensembles
Adria Ruiz, Jakob Verbeek
摘要
Inference in deep neural networks can be computationally expensive, and networks capable of anytime inference are important in scenarios where the amount of compute or quantity of input data varies over time. In such networks the inference process can interrupted to provide a result faster, or continued to obtain a more accurate result. We propose Hierarchical Neural Ensembles (HNE), a novel framework to embed an ensemble of multiple networks in a hierarchical tree structure, sharing intermediate layers. In HNE we control the complexity of inference on-the-fly by evaluating more or less models in the ensemble. Our second contribution is a novel hierarchical distillation method to boost the prediction accuracy of small ensembles. This approach leverages the nested structure of our ensembles, to optimally allocate accuracy and diversity across the individual models. Our experiments show that, compared to previous anytime inference models, HNE provides state-of-the-art accuracy-computate trade-offs on the CIFAR-10/100 and ImageNet datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Self-Contrastive Learning: Single-Viewed Supervised Contrastive Framework Using Sub-networkSangmin Bae, Sungnyun Kim, Jongwoo Ko, Gihun Lee 等AAAI 2023 · 被引用 14 次
- ProGMLP: A Progressive Framework for GNN-to-MLP Knowledge Distillation with Efficient Trade-offsWeigang Lu, Ziyu Guan, Wei Zhao, Yaming Yang 等AAAI 2026 · 被引用 1 次
- TIPS: Topologically Important Path Sampling for Anytime Neural NetworksGuihong Li, Kartikeya Bhardwaj, Yuedong Yang, Radu MarculescuICML 2023
- Neural Parameter Allocation SearchBryan A. Plummer, Nikoli Dryden, Julius Frost, Torsten Hoefler 等ICLR 2022
- Progressive Ensemble Distillation: Building Ensembles for Efficient InferenceDon Kurian Dennis, Abhishek Shetty, Anish Prasad Sevekari, Kazuhito Koishida 等NeurIPS 2023
它引用的顶会 Paper8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 被引用 273 次
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
相关 Paper
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen 等AAAI 2023 · 被引用 16 次
- AnyDA: Anytime Domain AdaptationOmprakash Chakraborty, Aadarsh Sahoo, Rameswar Panda, Abir DasICLR 2023
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 被引用 10 次
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 被引用 5 次
