Taming Dynamism on GPUs: Cross-SM Kernel Fusion via SM Cooperation and Just-in-Time Reduction
Jingkai He, Guangda Sun, TianJian Li, Dong Du, Yubin Xia, Haibo Chen
2026Year
Abstract
Modern LLM architectures and systems render GPU kernel input shapes increasingly dynamic, e.g., conditional expert routing in MoE and batches with variable-length sequences. This dynamism degrades performance in both expert-tuned kernels and compiler frameworks. We identify the root cause as dynamic cross-SM data dependencies across execution phases, which force distinct phases into separate kernels for correctness. The resulting kernel boundaries amplify intra-phase workload imbalance across SMs, and compel intermediate states to expensive round trips through off-chip memory.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- LLM-42: Enabling Determinism in LLM Inference with Verified SpeculationRaja Gond, Aditya K Kamath, Ramachandran Ramjee, Ashish PanwarSOSP 2026
- SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K RoutingZewen Jin, Shen Fu, Chengjie Tang, Youhui Bai et al.AAAI 2026
- Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUsBojian Zheng, Cody Hao Yu, Jie Wang, Yaoyao Ding et al.MICRO 2023 · 4 citations
- Relax: Composable Abstractions for End-to-End Dynamic Machine LearningRuihang Lai, Junru Shao, Siyuan Feng, Steven Lyubomirsky et al.ASPLOS 2025 · 15 citations
- Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language ModelsYapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang et al.EuroSys 2026
