Leveraging Domain Information for the Efficient Automated Design of Deep Learning Accelerators
Chirag Sakhuja, Zhan Shi, Calvin Lin
摘要
Deep learning accelerators are important tools for feeding the growing demand for deep learning applications. The automated design of such accelerators-which is important for reducing development costs-can be viewed as a search over a vast and complex design space that consists of all possible accelerators and all the possible software that could run on them.
Unfortunately, this search is complicated by the existence of many ordinal and categorical values, which are critical to explore for the ultimate design but are not handled well by existing search techniques.
This paper presents a technique for efficiently searching this space by injecting domain information-in this case information about hardware/software (HW/SW) co-design-into the automated search process. Specifically, this paper introduces a novel Bayesian optimization framework called daBO (domain-aware BO) that accepts domain information as input, including those describing ordinal and categorical values.
This paper also introduces Spotlight, a design tool based on daBO, and this paper empirically shows that Spotlight produces accelerator designs and software schedules that are orders of magnitude better than those created by the state-of-the-art. For example, for the ResNet-50 deep learning model, Spotlight produces a HW/SW configuration that reduces delay by 135× over the configuration produced by ConfuciuX, a state-of-theart HW/SW co-design tool, and Spotlight reduces energy-delay product (EDP) by 44× over an Eyeriss-like accelerator, which is an edge-scale hand-designed accelerator. In the realm of cloud-scale accelerators, Spotlight reduces the EDP of a scaledup Eyeriss-like accelerator by 23×. Our evaluation shows that Spotlight benefits from the efficiency of daBO, which allows Spotlight to identify accelerator designs and software schedules that prior work cannot identify.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DOSA: Differentiable Model-Based One-Loop Search for DNN AcceleratorsCharles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar 等MICRO 2023 · 被引用 19 次
- CATransformers: Carbon Aware Transformers Through Joint Model-Hardware OptimizationIrene Wang, Mostafa Elhoushi, Ekin Sumbul, Samuel Hsia 等NeurIPS 2025 · 被引用 8 次
- Integrated Hardware Architecture and Device Placement SearchIrene Wang, Jakub Tarnawski, Amar Phanishayee, Divya MahajanICML 2024 · 被引用 4 次
它引用的顶会 Paper12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter 等ASPLOS 2020 · 被引用 237 次
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksLei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon 等DAC 2020 · 被引用 115 次
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement LearningSheng-Chun Kao, Geonhwa Jeong, Tushar KrishnaMICRO 2020 · 被引用 115 次
相关 Paper
- UNICO: Unified Hardware Software Co-Optimization for Robust Neural Network AccelerationBahador Rashidi, Chao Gao, Shan Lu, Zhisheng Wang 等MICRO 2023 · 被引用 6 次
- Explainable-DSE: An Agile and Explainable Exploration of Efficient HW/SW Codesigns of Deep Learning Accelerators Using Bottleneck AnalysisShail Dave, Tony Nowatzki, Aviral ShrivastavaASPLOS 2023 · 被引用 7 次
- HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationQingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu 等ISCA 2021 · 被引用 73 次
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu 等ASPLOS 2022 · 被引用 48 次
- Glimpse: mathematical embedding of hardware specification for neural compilationByung Hoon Ahn, Sean Kinzer, Hadi EsmaeilzadehDAC 2022 · 被引用 4 次
