Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture Search
Pengtao Xie, Xuefeng Du
摘要
Knowledge distillation has shown great effectiveness for improving neural architecture search (NAS). Mutual knowledge distillation (MKD), where a group of models mutually generate knowledge to train each other, has achieved promising results in many applications. In existing MKD methods, mutual knowledge distillation is performed between models without scrutiny: a worse-performing model is allowed to generate knowledge to train a better-performing model, which may lead to collective failures. To address this problem, we propose a performance-aware MKD (PAMKD) approach for NAS, where knowledge generated by model <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> is allowed to train model <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> only if the performance of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> is better than B. We propose a three-level optimization framework to formulate PAMKD, where three learning stages are performed end-to-end: 1) each model trains an initial model independently; 2) the initial models are evaluated on a validation set and better-performing models generate knowledge to train worse-performing models; 3) architectures are updated by minimizing a validation loss. Experimental results on a variety of datasets demonstrate that our method is effective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Betty: An Automatic Differentiation Library for Multilevel OptimizationSang Keun Choe, Willie Neiswanger, Pengtao Xie, Eric P. XingICLR 2023 · 被引用 5 次
- Long-Tailed Question Answering in an Open WorldYi Dai, Hao Lang, Yinhe Zheng, Fei Huang 等ACL 2023 · 被引用 3 次
- Improving Bi-level Optimization Based Methods with Inspiration from Humans' Classroom Study TechniquesPengtao XieICML 2023 · 被引用 1 次
- FLAME: Condensing Ensemble Diversity into a Single Network for Efficient Sequential RecommendationWooJoo Kim, JunYoung Kim, Jaehyung Lim, SeongJin Choi 等SIGIR 2026
它引用的顶会 Paper26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 被引用 825 次
- Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationXin Chen, Lingxi Xie, Jun Wu, Qi TianICCV 2019 · 被引用 725 次
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen 等ICLR 2020 · 被引用 691 次
相关 Paper
- Search to Distill: Pearls Are Everywhere but Not the EyesYu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli 等CVPR 2020
- Improving Differentiable Neural Architecture Search by Encouraging TransferabilityParth Sheth, Pengtao XieICLR 2023
- Meta-prediction Model for Distillation-Aware NAS on Unseen DatasetsHayeon Lee, Sohyun An, Minseon Kim, Sung Ju HwangICLR 2023 · 被引用 2 次
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 被引用 48 次
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey 等NeurIPS 2022 · 被引用 21 次
