Heron: Automatically Constrained High-Performance Library Generation for Deep Learning Accelerators
Jun Bi, Qi Guo, Xiaqing Li, Yongwei Zhao, Yuanbo Wen, Yuxuan Guo, Enshuai Zhou, Xing Hu, Zidong Du, Ling Li, Huaping Chen, Tianshi Chen
摘要
Deep Learning Accelerators (DLAs) are effective to improve both performance and energy efficiency of compute-intensive deep learning algorithms. A flexible and portable mean to exploit DLAs is using high-performance software libraries with well-established APIs, which are typically either manually implemented or automatically generated by exploration-based compilation approaches. Though exploration-based approaches significantly reduce programming efforts, they fail to find optimal or near-optimal programs from a large but low-quality search space because the massive inherent constraints of DLAs cannot be accurately characterized.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program TuningLiang Qiao, Jun Shi, Xiaoyu Hao, Xi Fang 等ASPLOS 2025 · 被引用 5 次
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
- LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control FlowKaiyan Chang, Wenlong Zhu, Shengwen Liang, Huawei Li 等MICRO 2025 · 被引用 1 次
相关 Paper
- Felix: Optimizing Tensor Programs with Gradient DescentYifan Zhao, Hashim Sharif, Vikram S. Adve, Sasa MisailovicASPLOS 2024 · 被引用 13 次
- Explainable-DSE: An Agile and Explainable Exploration of Efficient HW/SW Codesigns of Deep Learning Accelerators Using Bottleneck AnalysisShail Dave, Tony Nowatzki, Aviral ShrivastavaASPLOS 2023 · 被引用 7 次
- Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex TessellationJianxing Xu, Yuanbo Wen, Zikang Liu, Ruibai Xu 等ASPLOS 2025 · 被引用 2 次
- HARP: holistic analysis for refactoring Python-based analytics programsWeijie Zhou, Yue Zhao, Guoqiang Zhang, Xipeng ShenICSE 2020 · 被引用 10 次
- A comprehensive study of deep learning compiler bugsQingchao Shen, Haoyang Ma, Junjie Chen, Yongqiang Tian 等FSE 2021 · 被引用 123 次
