NASA: Accelerating Neural Network Design with a NAS Processor
Xiaohan Ma, Chang Si, Ying Wang, Cheng Liu, Lei Zhang
Abstract
Neural network search (NAS) projects a promising direction to automate the design process of efficient and powerful neural network architectures. Nevertheless, the NAS techniques have to dynamically generate a large number of candidate neural networks, and iteratively train and evaluate these on-line generated network architectures, thus they are extremely time-consuming even when deployed on large GPU clusters, which dramatically hinders the adoption of NAS. Though recently there are many specialized architectures proposed to accelerate the training or inference of neural networks, we observe that existing neural network accelerators are typically targeted at static neural network architectures, and they are not suitable to accelerate the evaluation of the dynamical neural network candidates evolving during the NAS process, which cannot be deployed onto current accelerators via the off-line compilation.To enable rapid and energy-efficient NAS in compact single-chip solutions, we propose NASA, a specialized architecture for one-shot based NAS acceleration. It is able to generate, schedule, and evaluate the candidate neural network architectures for the target machine learning workload with high speed, significantly alleviating the processing bottleneck of one-shot NAS. Motivated by the observation that there are considerable computation sharing opportunities among the different neural network candidates generated in one-shot NAS, NASA is equipped with an on-chip network fusion unit to remove the redundant computation during the network mapping stage. In addition, the NASA accelerator can partition and re-schedule the candidate neural network architectures at fine-granularity to maximize the chance of data reuse and improve the utilization of the accelerator arrays integrated to accelerate network evaluation. According to our experiments on multiple one-shot NAS tasks, NASA achieves 33.52× performance speedup and 214.33× energy consumption reduction on average when compared to aCPU-GPU system.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b35d96f5-a14c-409d-b60d-e1da9db8337cCited by top-tier papers2
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsJingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei et al.HPCA 2024 · 65 citations
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 11 citations
Related papers
- NAS-SE: Designing A Highly-Efficient In-Situ Neural Architecture Search Engine for Large-Scale DeploymentQiyu Wan, Lening Wang, Jing Wang, Shuaiwen Leon Song et al.MICRO 2023 · 2 citations
- Searching by Generating: Flexible and Efficient One-Shot NAS With Architecture GeneratorSian-Yao Huang, Wei-Ta ChuCVPR 2021
- Few-Shot Neural Architecture SearchYiyang Zhao, Linnan Wang, Yuandong Tian, Rodrigo Fonseca et al.ICML 2021 · 100 citations
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONASHan Shi, Renjie Pi, Hang Xu, Zhenguo Li et al.NeurIPS 2020 · 148 citations
