Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox Model
Dongdong Wang, Yandong Li, Liqiang Wang, Boqing Gong
Abstract
We study how to train a student deep neural network for visual recognition by distilling knowledge from a blackbox teacher model in a data-efficient manner. Progress on this problem can significantly reduce the dependence on large-scale datasets for learning high-performing visual recognition models. There are two major challenges. One is that the number of queries into the teacher model should be minimized to save computational and/or financial costs. The other is that the number of images used for the knowledge distillation should be small; otherwise, it violates our expectation of reducing the dependence on large-scale datasets. To tackle these challenges, we propose an approach that blends mixup and active learning. The former effectively augments the few unlabeled images by a big pool of synthetic images sampled from the convex hull of the original images, and the latter actively chooses from the pool hard examples for the student neural network and query their labels from the teacher model. We validate our approach with extensive experiments. 1 . 1 Code and models: https://github.com/dwang181/active-mixup * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05ce9bcf-8121-4c94-84dd-db459d9bd0ffCited by top-tier papers10
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
- CTIN: Robust Contextual Transformer Network for Inertial NavigationBingbing Rao, Ehsan Kazemi, Yifan Ding, Devu M. Shila et al.AAAI 2022 · 66 citations
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box ModelZi WangICML 2021 · 56 citations
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song et al.AAAI 2021 · 55 citations
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 5 citations
Builds on14
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Learning Lightweight Lane Detection CNNs by Self Attention DistillationYuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change LoyICCV 2019 · 666 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
Related papers
- IDEAL: Query-Efficient Data-Free Learning from Black-Box ModelsJie Zhang, Chen Chen, Lingjuan LyuICLR 2023 · 5 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 2 citations
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li et al.NeurIPS 2021 · 22 citations
- Weighted Distillation with Unlabeled ExamplesFotis Iliopoulos, Vasilis Kontonis, Cenk Baykal, Gaurav Menghani et al.NeurIPS 2022 · 20 citations
