Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core Processor
Xinxin Qi, Jianbin Fang, Peng Zhang, Yonggang Che, Jie Ren
摘要
Optimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen 等HPCA 2023 · 被引用 46 次
- CoSA: Scheduling by Constrained Optimization for Spatial AcceleratorsQijing Huang, Aravind Kalaiah, Minwoo Kang, James Demmel 等ISCA 2021 · 被引用 120 次
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang 等DAC 2020 · 被引用 15 次
- Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesZelin Du, Wei Zhang, Zimeng Zhou, Zili Shao 等DAC 2023 · 被引用 9 次
- Deinsum: Practically I/O Optimal Multi-Linear AlgebraAlexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun, Timo Schneider 等SC 2022 · 被引用 2 次
