Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion
Size Zheng, Siyuan Chen, Peidi Song, Renze Chen, Xiuhong Li, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang
Abstract
Machine learning models with various tensor operators are becoming ubiquitous in recent years. There are two types of operators in machine learning: compute-intensive operators (e.g., GEMM and convolution) and memory-intensive operators (e.g., ReLU and softmax). In emerging machine learning models, compute-intensive operators are usually organized in a chain structure. With the continual specialization of hardware, the gap between computing performance and memory bandwidth has become more prominent. Consequently, the implementations of many compute-intensive operator chains are bounded by memory bandwidth, and generating fused kernels to improve locality for these compute-intensive operators becomes necessary. But in existing machine learning compilers, there lack both precise analysis and efficient optimization for compute-intensive operator chains on different accelerators. As a result, they usually produce sub-optimal performance for these operator chains.
In this paper, we propose Chimera, an optimizing framework that can efficiently improve the locality of compute-intensive operator chains on different hardware accelerators. In Chimera, each compute-intensive operator is composed of a series of computation blocks. To generate efficient fused kernels for the operator chains, optimizations for both inter-block and intrablock are required. For inter-block optimization, Chimera decides the optimized block execution order by minimizing the data movement volume among blocks using an analytical model. For intra-block optimization, Chimera uses unified replaceable micro kernels to apply hardware-specific optimizations for different accelerators. Finally, Chimera generates fused kernels for computeintensive operator chains. Evaluation of batch GEMM chains and convolution chains on CPU, GPU, and NPU shows that Chimera achieves up to 2.87×, 2.29×, and 2.39× speedups to hand-tuned libraries. Compared to state-of-the-art compilers, the speedups are up to 2.29×, 1.64×, and 1.14× for CPU, GPU, and NPU. TABLE II THE COMPARISON OF PREVIOUS REPRESENTATIVE WORK AND CHIMERA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c863fc7c-a63a-413d-a395-8b0d4611bda9Cited by top-tier papers11
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu et al.NeurIPS 2024 · 56 citations
- TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based AnalysisSize Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia et al.MICRO 2023 · 31 citations
- M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeWeiming Hu, Haoyan Zhang, Cong Guo, Yu Feng et al.HPCA 2025 · 19 citations
- MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNNRenze Chen, Zijian Ding, Size Zheng, Chengrui Zhang et al.ASPLOS 2024 · 14 citations
- Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoCSize Zheng, Siyuan Chen, Yun LiangDAC 2023 · 10 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter et al.ASPLOS 2020 · 237 citations
Related papers
- Chimera: Communication Fusion for Hybrid Parallelism in Large Language ModelsLe Qin, Junwei Cui, Weilin Cai, Jiayi HuangISCA 2025 · 11 citations
- AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesZhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long et al.ASPLOS 2022 · 78 citations
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 9 citations
- OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear ProgrammingJiantao Liu, Minxuan Zhou, Yue Pan, Chien-Yi Yang et al.ISCA 2025 · 6 citations
- Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core ProcessorXinxin Qi, Jianbin Fang, Peng Zhang, Yonggang Che et al.SC 2025 · 4 citations
