Taming Dynamism on GPUs: Cross-SM Kernel Fusion via SM Cooperation and Just-in-Time Reduction
Jingkai He, Guangda Sun, TianJian Li, Dong Du, Yubin Xia, Haibo Chen
2026年份
摘要
Modern LLM architectures and systems render GPU kernel input shapes increasingly dynamic, e.g., conditional expert routing in MoE and batches with variable-length sequences. This dynamism degrades performance in both expert-tuned kernels and compiler frameworks. We identify the root cause as dynamic cross-SM data dependencies across execution phases, which force distinct phases into separate kernels for correctness. The resulting kernel boundaries amplify intra-phase workload imbalance across SMs, and compel intermediate states to expensive round trips through off-chip memory.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LLM-42: Enabling Determinism in LLM Inference with Verified SpeculationRaja Gond, Aditya K Kamath, Ramachandran Ramjee, Ashish PanwarSOSP 2026
- SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K RoutingZewen Jin, Shen Fu, Chengjie Tang, Youhui Bai 等AAAI 2026
- Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUsBojian Zheng, Cody Hao Yu, Jie Wang, Yaoyao Ding 等MICRO 2023 · 被引用 4 次
- Relax: Composable Abstractions for End-to-End Dynamic Machine LearningRuihang Lai, Junru Shao, Siyuan Feng, Steven Lyubomirsky 等ASPLOS 2025 · 被引用 15 次
- Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language ModelsYapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang 等EuroSys 2026
