M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification
Mengting He, Shihao Xia, Haomin Jia, Wenfei Wu, Linhai Song
摘要
Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is implicit and poorly specified. Models and kernels evolve independently and often make incompatible assumptions about tensor shapes and input sizes, leading to subtle memory bugs in CUDA kernels. These bugs can crash inference services, corrupt model weights, or be exploited by remote adversaries. Existing techniques either incur prohibitive runtime overhead, require specialized hardware, or fail to handle dynamic tensor shapes and variable kernel launch configurations, leaving the CUDA memory bugs largely unaddressed.
This paper presents M2K, a fully automated framework that makes the model-kernel interface explicit and leverages it to detect memory bugs in CUDA kernels used in LLM inference systems. M2K consists of two components. HFProbe traces model execution without GPU hardware, classifies kernel arguments into model-fixed and user-variable, and emits symbolic constraints that capture the interface. cuK-LEE then performs symbolic execution on CUDA kernels to pinpoint memory bugs under the interface constraints, modeling tensors as disjoint memory regions and treating thread identifiers symbolically to scale to thousands of threads. In the evaluation, M2K discovers 181 previously unknown bugs in real LLM inference systems, while producing only nine false positives, demonstrating its effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Simulee: detecting CUDA synchronization bugs via memory-access modelingMingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang 等ICSE 2020 · 被引用 26 次
- CuSafe: Capturing Memory Corruption on NVIDIA GPUsHongyi Lu, Fengwei Zhang, Zhenkai Zhang, Shuai Wang 等USENIX Security 2026
- Modeling and analyzing evaluation cost of CUDA kernelsStefan K. Muller, Jan HoffmannPOPL 2021 · 被引用 15 次
- From Prompt to Pwn: Exploiting GPU Memory Errors During ML InferenceJonas Roels, Adriaan Jacobs, Silviu Vlasceanu, Mahmoud Ammar 等CCS 2026
- AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsShiheng Cao, Junmin Wu, Junshi Chen, Hong An 等ISCA 2025 · 被引用 5 次
