SC2025Top-tier venue
Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core Processor
Xinxin Qi, Jianbin Fang, Peng Zhang, Yonggang Che, Jie Ren
Abstract
Optimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 10b99935-066e-4bb9-98e5-46137bd322cdCited by top-tier papers1
Ask how each one uses itRelated papers
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen et al.HPCA 2023 · 46 citations
- CoSA: Scheduling by Constrained Optimization for Spatial AcceleratorsQijing Huang, Aravind Kalaiah, Minwoo Kang, James Demmel et al.ISCA 2021 · 120 citations
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang et al.DAC 2020 · 15 citations
- Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesZelin Du, Wei Zhang, Zimeng Zhou, Zili Shao et al.DAC 2023 · 9 citations
- Deinsum: Practically I/O Optimal Multi-Linear AlgebraAlexandros Nikolaos Ziogas, Grzegorz Kwasniewski, Tal Ben-Nun, Timo Schneider et al.SC 2022 · 2 citations
