Spying on the Floating Point Behavior of Existing, Unmodified Scientific Applications
Peter A. Dinda, Alex Bernat, Conor Hetland
摘要
Scientific (and other) applications are critically dependent on calculations done using IEEE floating point arithmetic. A number of concerns have been raised about correctness in such applications given the numerous gotchas the IEEE standard presents for developers, as well as the complexity of its implementation at the hardware and compiler levels. The standard and its implementations do provide mechanisms for analyzing floating point arithmetic as it executes, making it possible to find and track problematic operations. However, this capability is seldom used in practice. In response, we have developed FPSpy, a tool that provides this capability when operating underneath existing, unmodified x64 application binaries on Linux, including those using thread- and process-level parallelism. FPSpy can observe application behavior without any cooperation from the application or developer, and can potentially be deployed as part of a job launch process. We present the design, implementation, and performance evaluation of FPSpy. FPSpy operates conservatively, getting out of the way if the application itself begins to use any of the OS or hardware features that FPSpy depends on. Its overhead can be throttled, allowing a tradeoff between which and how many unusual events are to be captured, and the slowdown incurred by the application, with the low point providing virtually zero slowdown. We evaluated FPSpy by using it to methodically study seven widely-used applications/frameworks from a range of domains (five of which are in the NSF XSEDE top-20), as well as the NAS and PARSEC benchmark suites. All told, these comprise about 7.5 million lines of source code in a wide range of languages, and parallelism models (including OpenMP and MPI). FPSpy was able to produce trace information for all of them. The traces show that problematic floating point events occur in both the applications and the benchmarks. Analysis of the rounding behavior captured in our traces also suggests the feasibility of an approach to adding adaptive precision underneath existing, unmodified binaries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Parallel shadow execution to accelerate the debugging of numerical errorsSangeeta Chowdhary, Santosh NagarakatteFSE 2021 · 被引用 21 次
- Fast shadow execution for debugging numerical errors using error free transformationsSangeeta Chowdhary, Santosh NagarakatteOOPSLA 2022 · 被引用 13 次
- Design and Evaluation of GPU-FPX: A Low-Overhead tool for Floating-Point Exception Detection in NVIDIA GPUsXinyi Li, Ignacio Laguna, Bo Fang, Katarzyna Swirydowicz 等HPDC 2023 · 被引用 12 次
- Finding Inputs that Trigger Floating-Point Exceptions in GPUs via Bayesian OptimizationIgnacio Laguna, Ganesh GopalakrishnanSC 2022 · 被引用 10 次
- RAPTOR: Practical Numerical Profiling of Scientific ApplicationsFaveo Hoerold, Ivan R. Ivanov, Akash Dhruv, William S. Moses 等SC 2025 · 被引用 4 次
相关 Paper
- Rigorous Floating-Point Round-Off Error Analysis in PRECiSA 4.0Laura Titolo, Mariano M. Moscato, Marco A. Feliú, Paolo Masci 等FM 2024 · 被引用 1 次
- FPVM: Towards a Floating Point Virtual MachinePeter A. Dinda, Nick Wanninger, Jiacheng Ma, Alex Bernat 等HPDC 2022 · 被引用 4 次
- FPBOXer: Efficient Input-Generation for Targeting Floating-Point Exceptions in GPU ProgramsAnh Tran, Ignacio Laguna, Ganesh GopalakrishnanHPDC 2024 · 被引用 3 次
- pLiner: isolating lines of floating-point code for compiler-induced variabilityHui Guo, Ignacio Laguna, Cindy Rubio-GonzálezSC 2020 · 被引用 15 次
- FloatGuard: Efficient Whole-Program Detection of Floating-Point Exceptions in AMD GPUsDolores Miao, Ignacio Laguna, Cindy Rubio-GonzálezHPDC 2025 · 被引用 2 次
