Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces
Bert Moons, Parham Noorzad, Andrii Skliar, Giovanni Mariani, Dushyant Mehta, Chris Lott, Tijmen Blankevoort
摘要
Current state-of-the-art Neural Architecture Search (NAS) methods neither efficiently scale to multiple hardware platforms, nor handle diverse architectural search-spaces. To remedy this, we present DONNA (Distilling Optimal Neural Network Architectures), a novel pipeline for rapid, scalable and diverse NAS, that scales to many user scenarios. DONNA consists of three phases. First, an accuracy predictor is built using blockwise knowledge distillation from a reference model. This predictor enables searching across diverse networks with varying macro-architectural parameters such as layer types and attention mechanisms, as well as across micro-architectural parameters such as block repeats and expansion rates. Second, a rapid evolutionary search finds a set of pareto-optimal architectures for any scenario using the accuracy predictor and on-device measurements. Third, optimal models are quickly fine-tuned to training-from-scratch accuracy. DONNA is up to 100× faster than MNasNet in finding state-of-the-art architectures on-device. Classifying ImageNet, DONNA architectures are 20% faster than EfficientNet-B0 and Mo-bileNetV2 on a Nvidia V100 GPU and 10% faster with 0.5% higher accuracy than MobileNetV2-1.4x on a Samsung S20 smartphone. In addition to NAS, DONNA is used for search-space extension and exploration, as well as hardware-aware model compression.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- BossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture SearchChanglin Li, Tao Tang, Guangrun Wang, Jiefeng Peng 等ICCV 2021 · 被引用 123 次
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu 等ICML 2021 · 被引用 52 次
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 被引用 36 次
- ZiCo: Zero-shot NAS via inverse Coefficient of Variation on GradientsGuihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu MarculescuICLR 2023 · 被引用 19 次
- LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture SearchChengquan Feng, Li Lyna Zhang, Yuanchi Liu, Jiahang Xu 等NSDI 2024 · 被引用 7 次
它引用的顶会 Paper9
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture SearchYuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen 等ICLR 2020 · 被引用 691 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
相关 Paper
- Block-Wisely Supervised Neural Architecture Search With Knowledge DistillationChanglin Li, Jiefeng Peng, Liuchun Yuan, Guangrun Wang 等CVPR 2020
- Search to Distill: Pearls Are Everywhere but Not the EyesYu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli 等CVPR 2020
- Meta-prediction Model for Distillation-Aware NAS on Unseen DatasetsHayeon Lee, Sohyun An, Minseon Kim, Sung Ju HwangICLR 2023 · 被引用 2 次
- FBNetV3: Joint Architecture-Recipe Search Using Predictor PretrainingXiaoliang Dai, Alvin Wan, Peizhao Zhang, Bichen Wu 等CVPR 2021
- MathNAS: If Blocks Have a Role in Mathematical Architecture DesignQinsi Wang, Jinghan Ke, Zhi Liang, Sihai ZhangNeurIPS 2023 · 被引用 6 次
