EDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI Solutions
Yuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu, Yao Chen, Jinjun Xiong, Wen-mei W. Hwu, Deming Chen
摘要
High quality AI solutions require joint optimization of AI algorithms and their hardware implementations. In this work, we are the first to propose a fully simultaneous, Efficient Differentiable DNN (deep neural network) architecture and implementation co-search (EDD) methodology. We formulate the co-search problem by fusing DNN search variables and hardware implementation variables into one solution space, and maximize both algorithm accuracy and hardware implementation quality. The formulation is differentiable with respect to the fused variables, so that gradient descent algorithm can be applied to greatly reduce the search time. The formulation is also applicable for various devices with different objectives. In the experiments, we demonstrate the effectiveness of our EDD methodology by searching for three representative DNNs, targeting low-latency GPU implementation and FPGA implementations with both recursive and pipelined architectures. Each model produced by EDD achieves similar accuracy as the best existing DNN models searched by neural architecture search (NAS) methods on ImageNet, but with superior performance obtained within 12 GPU-hour searches. Our DNN targeting GPU is 1.40× faster than the state-of-the-art solution reported in Proxyless [1], and our DNN targeting FPGA delivers 1.45× higher throughput than the state-of-the-art solution reported in DNNBuilder [2].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- NAAS: Neural Accelerator Architecture SearchYujun Lin, Mengtian Yang, Song HanDAC 2021 · 被引用 60 次
- DANCE: Differentiable Accelerator/Network Co-ExplorationKanghyun Choi, Deokki Hong, Hojae Yoon, Joonsang Yu 等DAC 2021 · 被引用 49 次
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu 等ASPLOS 2022 · 被引用 48 次
- Generic Neural Architecture Search via RegressionYuhong Li, Cong Hao, Pan Li, Jinjun Xiong 等NeurIPS 2021 · 被引用 40 次
- FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN TrainingYonggan Fu, Haoran You, Yang Zhao, Yue Wang 等NeurIPS 2020 · 被引用 36 次
相关 Paper
- Enabling hard constraints in differentiable neural network and accelerator co-explorationDeokki Hong, Kanghyun Choi, Hyeyoon Lee, Joonsang Yu 等DAC 2022 · 被引用 4 次
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and AcceleratorsYonggan Fu, Yongan Zhang, Yang Zhang, David D. Cox 等ICML 2021 · 被引用 23 次
- UniCoS: A Unified Neural and Accelerator Co-Search Framework for CNNs and ViTsWei Fu, Wenqi Lou, Cheng Tang, Hongbing Wen 等DAC 2025 · 被引用 1 次
- XPert: Peripheral Circuit & Neural Architecture Co-search for Area and Energy-efficient Xbar-based ComputingAbhishek Moitra, Abhiroop Bhattacharjee, Youngeun Kim, Priyadarshini PandaDAC 2023 · 被引用 8 次
- DOSA: Differentiable Model-Based One-Loop Search for DNN AcceleratorsCharles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar 等MICRO 2023 · 被引用 19 次
