Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators
Chirag Sakhuja, Zhan Shi, Calvin Lin
Abstract
Deep learning accelerators are important tools for feeding the growing demand for deep learning applications. The automated design of such accelerators-which is important for reducing development costs-can be viewed as a search over a vast and complex design space that consists of all possible accelerators and all the possible software that could run on them.
Unfortunately, this search is complicated by the existence of many ordinal and categorical values, which are critical to explore for the ultimate design but are not handled well by existing search techniques.
This paper presents a technique for efficiently searching this space by injecting domain information-in this case information about hardware/software (HW/SW) co-design-into the automated search process. Specifically, this paper introduces a novel Bayesian optimization framework called daBO (domain-aware BO) that accepts domain information as input, including those describing ordinal and categorical values.
This paper also introduces Spotlight, a design tool based on daBO, and this paper empirically shows that Spotlight produces accelerator designs and software schedules that are orders of magnitude better than those created by the state-of-the-art. For example, for the ResNet-50 deep learning model, Spotlight produces a HW/SW configuration that reduces delay by 135× over the configuration produced by ConfuciuX, a state-of-theart HW/SW co-design tool, and Spotlight reduces energy-delay product (EDP) by 44× over an Eyeriss-like accelerator, which is an edge-scale hand-designed accelerator. In the realm of cloud-scale accelerators, Spotlight reduces the EDP of a scaledup Eyeriss-like accelerator by 23×. Our evaluation shows that Spotlight benefits from the efficiency of daBO, which allows Spotlight to identify accelerator designs and software schedules that prior work cannot identify.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85e4f69b-94c6-4ec2-b757-b6f012c32203Cited by top-tier papers3
- DOSA: Differentiable Model-Based One-Loop Search for DNN AcceleratorsCharles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar et al.MICRO 2023 · 19 citations
- CATransformers: Carbon Aware Transformers Through Joint Model-Hardware OptimizationIrene Wang, Mostafa Elhoushi, Ekin Sumbul, Samuel Hsia et al.NeurIPS 2025 · 8 citations
- Integrated Hardware Architecture and Device Placement SearchIrene Wang, Jakub Tarnawski, Amar Phanishayee, Divya MahajanICML 2024 · 4 citations
Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter et al.ASPLOS 2020 · 237 citations
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksLei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon et al.DAC 2020 · 115 citations
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement LearningSheng-Chun Kao, Geonhwa Jeong, Tushar KrishnaMICRO 2020 · 115 citations
Related papers
- UNICO: Unified Hardware Software Co-Optimization for Robust Neural Network AccelerationBahador Rashidi, Chao Gao, Shan Lu, Zhisheng Wang et al.MICRO 2023 · 6 citations
- Explainable-DSE: An Agile and Explainable Exploration of Efficient HW/SW Codesigns of Deep Learning Accelerators Using Bottleneck AnalysisShail Dave, Tony Nowatzki, Aviral ShrivastavaASPLOS 2023 · 7 citations
- HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationQingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu et al.ISCA 2021 · 73 citations
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu et al.ASPLOS 2022 · 48 citations
- Glimpse: mathematical embedding of hardware specification for neural compilationByung Hoon Ahn, Sean Kinzer, Hadi EsmaeilzadehDAC 2022 · 4 citations
