cuCatch: A Debugging Tool for Efficiently Catching Memory Safety Violations in CUDA Applications
Mohamed Tarek Ibn Ziad, Sana Damani, Aamer Jaleel, Stephen W. Keckler, Mark Stephenson
摘要
MARK STEPHENSON, NVIDIA, USA CUDA, OpenCL, and OpenACC are the primary means of writing general-purpose software for NVIDIA GPUs, all of which are subject to the same well-documented memory safety vulnerabilities currently plaguing software written in C and C++. One can argue that the GPU execution environment makes software development more error prone. Unlike C and C++, CUDA features multiple, distinct memory spaces to map to the GPU's unique memory hierarchy, and a typical CUDA program has thousands of concurrently executing threads. Furthermore, the CUDA platform has fewer guardrails than CPU platforms that have been forced to incrementally adjust to a barrage of security attacks. Unfortunately, the peculiarities of the GPU make it difficult to directly port memory safety solutions from the CPU space.
This paper presents cuCatch, a new memory safety error detection tool designed specifically for the CUDA programming model. cuCatch combines optimized compiler instrumentation with driver support to implement a novel algorithm for catching spatial and temporal memory safety errors with low performance overheads. Our experimental results on a wide set of GPU applications show that cuCatch incurs a 19% runtime slowdown on average, which is orders of magnitude faster than state-of-the-art debugging tools on GPUs. Moreover, our quantitative evaluation demonstrates cuCatch's higher error detection coverage compared to prior memory safety tools. The combination of high error detection coverage and low runtime overheads makes cuCatch an ideal candidate for accelerating memory safety debugging for GPU applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout 等HPCA 2025 · 被引用 12 次
- GPU Memory Exploitation for Fun and ProfitYanan Guo, Zhenkai Zhang, Jun YangUSENIX Security 2024 · 被引用 12 次
- GPUBreach: Privilege Escalation Attacks on GPUs Using RowhammerChris S. Lin, Yuqin Yan, Guozhen Ding, Joyce Qu 等S&P 2026 · 被引用 8 次
- Adaptive CHERI Compartmentalization for Heterogeneous AcceleratorsJianyi Cheng, A. Theodore Markettos, Alexandre Joannou, Paul Metzger 等ISCA 2025 · 被引用 4 次
- Let-Me-In: (Still) Employing In-pointer Bounds Metadata for Fine-grained GPU Memory SafetyJaewon Lee, Euijun Chung, Saurabh Singh, Seonjin Na 等HPCA 2025 · 被引用 3 次
它引用的顶会 Paper7
- SoK: Sanitizing for SecurityDokyung Song, Julian Lettner, Prabhu Rajasekaran, Yeoul Na 等S&P 2019 · 被引用 196 次
- Hardware-based Always-On Heap Memory SafetyYonghae Kim, Jaekyu Lee, Hyesoon KimMICRO 2020 · 被引用 41 次
- CHEx86: Context-Sensitive Enforcement of Memory Safety via Microcode-Enabled CapabilitiesRasool Sharifi, Ashish VenkatISCA 2020 · 被引用 28 次
- Simulee: detecting CUDA synchronization bugs via memory-access modelingMingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang 等ICSE 2020 · 被引用 26 次
- No-FAT: Architectural Support for Low Overhead Memory Safety ChecksMohamed Tarek Ibn Ziad, Miguel A. Arroyo, Evgeny Manzhosov, Ryan Piersma 等ISCA 2021 · 被引用 26 次
相关 Paper
- CuSafe: Capturing Memory Corruption on NVIDIA GPUsHongyi Lu, Fengwei Zhang, Zhenkai Zhang, Shuai Wang 等USENIX Security 2026
- SuperCollider: Scalable and Effective Data Race Detection for CUDAMark Stephenson, Sana Damani, Mohamed Tarek Ibn Ziad, Anis Ladram 等PLDI 2026
- From Prompt to Pwn: Exploiting GPU Memory Errors During ML InferenceJonas Roels, Adriaan Jacobs, Silviu Vlasceanu, Mahmoud Ammar 等CCS 2026
- Securing GPU via region-based bounds checkingJaewon Lee, Yonghae Kim, Jiashen Cao, Euna Kim 等ISCA 2022 · 被引用 19 次
- sfGPUMC: A Stateless Model Checker for GPU Weak Memory ConcurrencySoham Chakraborty, S. Krishna, Andreas Pavlogiannis, Omkar TuppeCAV 2025 · 被引用 2 次
