EDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI Solutions
Yuhong Li, Cong Hao, Xiaofan Zhang, Xinheng Liu, Yao Chen, Jinjun Xiong, Wen-mei W. Hwu, Deming Chen
Abstract
High quality AI solutions require joint optimization of AI algorithms and their hardware implementations. In this work, we are the first to propose a fully simultaneous, Efficient Differentiable DNN (deep neural network) architecture and implementation co-search (EDD) methodology. We formulate the co-search problem by fusing DNN search variables and hardware implementation variables into one solution space, and maximize both algorithm accuracy and hardware implementation quality. The formulation is differentiable with respect to the fused variables, so that gradient descent algorithm can be applied to greatly reduce the search time. The formulation is also applicable for various devices with different objectives. In the experiments, we demonstrate the effectiveness of our EDD methodology by searching for three representative DNNs, targeting low-latency GPU implementation and FPGA implementations with both recursive and pipelined architectures. Each model produced by EDD achieves similar accuracy as the best existing DNN models searched by neural architecture search (NAS) methods on ImageNet, but with superior performance obtained within 12 GPU-hour searches. Our DNN targeting GPU is 1.40× faster than the state-of-the-art solution reported in Proxyless [1], and our DNN targeting FPGA delivers 1.45× higher throughput than the state-of-the-art solution reported in DNNBuilder [2].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 513c1d85-e995-4ae7-9467-3277c4c6b47dCited by top-tier papers13
- NAAS: Neural Accelerator Architecture SearchYujun Lin, Mengtian Yang, Song HanDAC 2021 · 60 citations
- DANCE: Differentiable Accelerator/Network Co-ExplorationKanghyun Choi, Deokki Hong, Hojae Yoon, Joonsang Yu et al.DAC 2021 · 49 citations
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu et al.ASPLOS 2022 · 48 citations
- Generic Neural Architecture Search via RegressionYuhong Li, Cong Hao, Pan Li, Jinjun Xiong et al.NeurIPS 2021 · 40 citations
- FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN TrainingYonggan Fu, Haoran You, Yang Zhao, Yue Wang et al.NeurIPS 2020 · 36 citations
Related papers
- Enabling hard constraints in differentiable neural network and accelerator co-explorationDeokki Hong, Kanghyun Choi, Hyeyoon Lee, Joonsang Yu et al.DAC 2022 · 4 citations
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and AcceleratorsYonggan Fu, Yongan Zhang, Yang Zhang, David D. Cox et al.ICML 2021 · 23 citations
- UniCoS: A Unified Neural and Accelerator Co-Search Framework for CNNs and ViTsWei Fu, Wenqi Lou, Cheng Tang, Hongbing Wen et al.DAC 2025 · 1 citation
- XPert: Peripheral Circuit & Neural Architecture Co-search for Area and Energy-efficient Xbar-based ComputingAbhishek Moitra, Abhiroop Bhattacharjee, Youngeun Kim, Priyadarshini PandaDAC 2023 · 8 citations
- DOSA: Differentiable Model-Based One-Loop Search for DNN AcceleratorsCharles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar et al.MICRO 2023 · 19 citations
