Progressive Ensemble Distillation: Building Ensembles for Efficient Inference
Don Kurian Dennis, Abhishek Shetty, Anish Prasad Sevekari, Kazuhito Koishida, Virginia Smith
摘要
We study the problem of progressive ensemble distillation: Given a large, pretrained teacher model , we seek to decompose the model into smaller, low-inference cost student models , such that progressively evaluating additional models in this ensemble leads to improved predictions. The resulting ensemble allows for flexibly tuning accuracy vs. inference cost at runtime, which is useful for a number of applications in on-device inference. The method we propose, B-DISTIL , relies on an algorithmic procedure that uses function composition over intermediate activations to construct expressive ensembles with similar performance as , but with smaller student models. We demonstrate the effectiveness of B-DISTIL by decomposing pretrained models across standard image, speech, and sensor datasets. We also provide theoretical guarantees in terms of convergence and generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ProGMLP: A Progressive Framework for GNN-to-MLP Knowledge Distillation with Efficient Trade-offsWeigang Lu, Ziyu Guan, Wei Zhao, Yaming Yang 等AAAI 2026 · 被引用 1 次
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
它引用的顶会 Paper7
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization SpaceArnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu 等CVPR 2022 · 被引用 64 次
相关 Paper
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 被引用 10 次
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 被引用 21 次
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon 等ICML 2025
- Reinforced Multi-Teacher Selection for Knowledge DistillationFei Yuan, Linjun Shou, Jian Pei, Wutao Lin 等AAAI 2021 · 被引用 155 次
- Boomerang Distillation Enables Zero-Shot Model Size InterpolationSara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero 等ICLR 2026 · 被引用 3 次
