NAS-SE: Designing A Highly-Efficient In-Situ Neural Architecture Search Engine for Large-Scale Deployment
Qiyu Wan, Lening Wang, Jing Wang, Shuaiwen Leon Song, Xin Fu
Abstract
The emergence of Neural Architecture Search (NAS) enables an automated neural network development process that potentially replaces manually-enabled machine learning expertise. A state-of-the-art NAS method, namely One-Shot NAS, has been proposed to drastically reduce the lengthy search time for a wide spectrum of conventional NAS methods. Nevertheless, the search cost is still prohibitively expensive for practical large-scale deployment with real-world applications. In this paper, we reveal that the fundamental cause for inefficient deployment of One-Shot NAS in both single-device and large-scale scenarios originates from the massive redundant off-chip weight access during the numerous DNN inference in sequential searching. Inspired by its algorithmic characteristics, we depart from the traditional CMOS-based architecture designs and propose a promising processing-in-memory design alternative to perform in-situ architecture search, which helps fundamentally address the redundancy issue. Moreover, we further discovered two major performance challenges of directly porting the searching process onto the existing PIM-based accelerators: severe pipeline contention and resource under-utilization. By leveraging these insights, we propose the first highly-efficient in-situ One-Shot NAS search engine design, named NAS-SE, for both single-device and large-scale deployment scenarios. NAS-SE is equipped with a two-phased network diversification strategy for eliminating resource contention, and a novel hardware mapping scheme for boosting the resource utilization by an order of magnitude. Our extensive evaluation demonstrates that NAS-SE significantly outperforms the state-of-the-art digital-based customized NAS accelerator (NASA) with an average speedup of 8.8 × and energy-efficiency improvement of 2.05 ×.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e49f24b-8725-4120-b0fb-1b59f8f0c309Related papers
- NASA: Accelerating Neural Network Design with a NAS ProcessorXiaohan Ma, Chang Si, Ying Wang, Cheng Liu et al.ISCA 2021 · 8 citations
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- Hyperscale Hardware Optimized Neural Architecture SearchSheng Li, Garrett Andersen, Tao Chen, Liqun Cheng et al.ASPLOS 2023 · 10 citations
- Searching by Generating: Flexible and Efficient One-Shot NAS With Architecture GeneratorSian-Yao Huang, Wei-Ta ChuCVPR 2021
- PreNAS: Preferred One-Shot Learning Towards Efficient Neural Architecture SearchHaibin Wang, Ce Ge, Hesen Chen, Xiuyu SunICML 2023 · 27 citations
