Finding Inputs that Trigger Floating-Point Exceptions in GPUs via Bayesian Optimization
Ignacio Laguna, Ganesh Gopalakrishnan
摘要
and can be tricky to debug, today developers lack practical solutions to detect or predict floating-point exceptions in GPUs. Most approaches on exception detection rely on hardware register flags [1], [2]; however, NVIDIA GPUs do not provide such register flags and CUDA provides no mechanism to detect such exceptions 1 . Compiler-based solutions, such as [3], require sources and can detect exceptions at runtime, but only for given inputs. Ideally, developers would want to know the inputs that induce exceptions so those inputs can be controlled and explored systematically during testing.
There is prior work on identifying inputs that induce floating-point exceptions in CPU programs [4]. While these methods could (in principle) be implemented on GPU kernels, their drawback is that they use SMT solvers and symbolic execution, which requires analyzing the source code. Unfortunately, practical GPU codes running on NVIDIA GPUs involve using proprietary accelerated libraries, such as cuBLAS, cuFFT, cuSOLVER, CUDA Math Lib, cuTENSOR, cuSPARSE, and cuDNN [5], for which the source code is not publicly available. Even popular machine learning frameworks, such as PyTorch [6], make heavy use of cuBLAS and cuDNN. As a result, methods that require the source code are limited to test accelerated libraries or code that use them.
Our Contributions. This paper presents Xscope 2 , a framework to find inputs that trigger floating-point exceptions in a GPU function f (x), where the function user has limited knowledge of how the function operates. More specifically, the source of f (x) is not available-the function is a black box from the user's perspective-and the user does not know a priori the input bounds that the function expects, i.e., inputs can be any normal floating-point number. Our method relies on using Bayesian optimization (BO) to explore f in a guided manner with the objective of pinpointing extreme cases in f . These extreme cases make f return the result of exceptions to the user, i.e., infinity (positive and negative), underflows (i.e., subnormal numbers), or NaN (not a number). Finally, the user is provided the inputs of f that triggered such exceptions. Developers can also use Xscope to test functions where the code is available and/or input bounds are known, which only 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Design and Evaluation of GPU-FPX: A Low-Overhead tool for Floating-Point Exception Detection in NVIDIA GPUsXinyi Li, Ignacio Laguna, Bo Fang, Katarzyna Swirydowicz 等HPDC 2023 · 被引用 12 次
- Characterizing Real-World Bugs in Tile Programs for Automated Bug DetectionRavishka Rathnasuriya, Zihe Song, Nidhi Majoju, Tingxi Li 等ISSTA 2026
它引用的顶会 Paper3
- Efficient generation of error-inducing floating-point inputs via symbolic executionHui Guo, Cindy Rubio-GonzálezICSE 2020 · 被引用 27 次
- Spying on the Floating Point Behavior of Existing, Unmodified Scientific ApplicationsPeter A. Dinda, Alex Bernat, Conor HetlandHPDC 2020 · 被引用 20 次
- pLiner: isolating lines of floating-point code for compiler-induced variabilityHui Guo, Ignacio Laguna, Cindy Rubio-GonzálezSC 2020 · 被引用 15 次
相关 Paper
- FPBOXer: Efficient Input-Generation for Targeting Floating-Point Exceptions in GPU ProgramsAnh Tran, Ignacio Laguna, Ganesh GopalakrishnanHPDC 2024 · 被引用 3 次
- FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUsDolores Miao, Ignacio Laguna, Cindy Rubio-GonzálezHPDC 2025 · 被引用 2 次
- An Investigation on Numerical Bugs in GPU Programs Towards Automated Bug DetectionRavishka Rathnasuriya, Nidhi Majoju, Zihe Song, Wei YangISSTA 2025
- Over-Synchronization in GPU ProgramsAjay Nayak, Arkaprava BasuMICRO 2024 · 被引用 2 次
- SuperCollider: Scalable and Effective Data Race Detection for CUDAMark Stephenson, Sana Damani, Mohamed Tarek Ibn Ziad, Anis Ladram 等PLDI 2026
