Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox Model
Dongdong Wang, Yandong Li, Liqiang Wang, Boqing Gong
摘要
We study how to train a student deep neural network for visual recognition by distilling knowledge from a blackbox teacher model in a data-efficient manner. Progress on this problem can significantly reduce the dependence on large-scale datasets for learning high-performing visual recognition models. There are two major challenges. One is that the number of queries into the teacher model should be minimized to save computational and/or financial costs. The other is that the number of images used for the knowledge distillation should be small; otherwise, it violates our expectation of reducing the dependence on large-scale datasets. To tackle these challenges, we propose an approach that blends mixup and active learning. The former effectively augments the few unlabeled images by a big pool of synthetic images sampled from the convex hull of the original images, and the latter actively chooses from the pool hard examples for the student neural network and query their labels from the teacher model. We validate our approach with extensive experiments. 1 . 1 Code and models: https://github.com/dwang181/active-mixup * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 等ICLR 2021 · 被引用 90 次
- CTIN: Robust Contextual Transformer Network for Inertial NavigationBingbing Rao, Ehsan Kazemi, Yifan Ding, Devu M. Shila 等AAAI 2022 · 被引用 66 次
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box ModelZi WangICML 2021 · 被引用 56 次
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song 等AAAI 2021 · 被引用 55 次
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 被引用 5 次
它引用的顶会 Paper14
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- Learning Lightweight Lane Detection CNNs by Self Attention DistillationYuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change LoyICCV 2019 · 被引用 666 次
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou 等ICCV 2019 · 被引用 211 次
相关 Paper
- IDEAL: Query-Efficient Data-Free Learning from Black-Box ModelsJie Zhang, Chen Chen, Lingjuan LyuICLR 2023 · 被引用 5 次
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin 等ICLR 2020 · 被引用 469 次
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 被引用 2 次
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li 等NeurIPS 2021 · 被引用 22 次
- Weighted Distillation with Unlabeled ExamplesFotis Iliopoulos, Vasilis Kontonis, Cenk Baykal, Gaurav Menghani 等NeurIPS 2022 · 被引用 20 次
