TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
Shixun Wu, Yujia Zhai, Jinyang Liu, Jiajun Huang, Zizhe Jian, Huangliang Dai, Sheng Di, Franck Cappello, Zizhong Chen
Abstract
GPU-based fast Fourier transform (FFT) is extremely important for scientific computing and signal processing. However, we find the inefficiency of existing FFT libraries and the absence of fault tolerance against soft error. To address these issues, we introduce TurboFFT, a new FFT prototype co-designed for high performance and online fault tolerance. For FFT, we propose an architecture-aware, padding-free, and template-based prototype to maximize hardware resource utilization, achieving a competitive or superior performance compared to the state-of-the-art closed-source library, cuFFT. For fault tolerance, we 1) explore algorithm-based fault tolerance (ABFT) at the thread and threadblock levels to reduce additional memory footprint, 2) address the error propagation by introducing a two-side ABFT with location encoding, and 3) further modify the threadblock-level FFT from 1-transaction to multi-transaction in order to bring more parallelism for ABFT. Our two-side strategy enables online correction without additional global memory while our multi-transaction design averages the expensive threadblock-level reduction in ABFT with zero additional operations. Experimental results on an NVIDIA A100 server GPU and a Tesla Turing T4 GPU demonstrate that TurboFFT without fault tolerance is comparable to or up to 300% faster than cuFFT and outperforms VkFFT. TurboFFT with fault tolerance maintains an overhead of 7% to 15%, even under tens of error injections per minute for both FP32 and FP64.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 949ba0f4-3a5c-41a6-8aa8-8d4414b02f00Cited by top-tier papers2
- FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsHairui Zhao, Qi Tian, Hongliang Li, Zizhong ChenUSENIX ATC 2025 · 6 citations
- TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPUShixun Wu, Yujia Zhai, Huangliang Dai, Yue Zhu et al.SC 2025 · 5 citations
Builds on2
- High-performance Effective Scientific Error-bounded Lossy Compression with Auto-tuned Multi-component InterpolationJinyang Liu, Sheng Di, Kai Zhao, Xin Liang et al.SIGMOD 2024 · 29 citations
- cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationJinyang Liu, Jiannan Tian, Shixun Wu, Sheng Di et al.SC 2024 · 17 citations
Related papers
- FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant AttentionHuangliang Dai, Shixun Wu, Jiajun Huang, Zizhe Jian et al.SC 2025 · 12 citations
- Arithmetic-intensity-guided fault tolerance for neural network inference on GPUsJack Kosaian, K. V. RashmiSC 2021 · 51 citations
- Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionLishan Yang, Bin Nie, Adwait Jog, Evgenia SmirniICSE 2021 · 25 citations
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong et al.ISCA 2026
- Asymmetric Resilience: Exploiting Task-Level Idempotency for Transient Error Recovery in Accelerator-Based SystemsJingwen Leng, Alper Buyuktosunoglu, Ramon Bertran, Pradip Bose et al.HPCA 2020 · 19 citations
