Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel Polymerization
Feng Yu, Guangli Li, Jiacheng Zhao, Huimin Cui, Xiaobing Feng, Jingling Xue
Abstract
In recent times, dynamic-shape neural networks have gained widespread usage in intelligent applications to address complex tasks, introducing challenges in optimizing tensor programs due to their dynamic nature. As the operators' shapes are determined at runtime in dynamic scenarios, the compilation process becomes expensive, limiting the practicality of existing static-shape tensor compilers. To address the need for effective and efficient optimization of dynamic-shape neural networks, this paper introduces MikPoly, a novel dynamic-shape tensor compiler based on micro-kernel polymerization. MikPoly employs a two-stage optimization approach, dynamically combining multiple statically generated micro-kernels using a lightweight cost model based on the shape of a tensor operator known at runtime. We evaluate the effectiveness of MikPoly by employing popular dynamic-shape operators and neural networks on two representative accelerators, namely GPU Tensor Cores and Ascend NPUs. Our experimental results demonstrate that MikPoly effectively optimizes dynamic-shape workloads, yielding an average performance improvement of 1.49× over state-of-the-art vendor libraries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef475f23-612f-4667-b34d-2c992356fa2eCited by top-tier papers1
Ask how each one uses itBuilds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
Related papers
- Relax: Composable Abstractions for End-to-End Dynamic Machine LearningRuihang Lai, Junru Shao, Siyuan Feng, Steven Lyubomirsky et al.ASPLOS 2025 · 15 citations
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu et al.SIGMOD 2024 · 14 citations
- SoD2: Statically Optimizing Dynamic Deep Neural Network ExecutionWei Niu, Gagan Agrawal, Bin RenASPLOS 2024 · 6 citations
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim et al.ASPLOS 2026 · 1 citation
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng et al.PLDI 2021 · 81 citations
