TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
Shixun Wu, Yujia Zhai, Huangliang Dai, Yue Zhu, Haiyang Hu, Zizhong Chen
摘要
Fourier Neural Operators (FNO) are widely used for learning partial differential equation solution operators. However, FNO lacks architecture-aware optimizations,with its Fourier layers executing FFT, filtering, GEMM, zero padding, and iFFT as separate stages, incurring multiple kernel launches and significant global memory traffic. We propose TurboFNO, the first fully fused FFT-GEMM-iFFT GPU kernel with built-in FFT optimizations. We first develop FFT and GEMM kernels from scratch, achieving performance comparable to or faster than the closed-source SOTA cuBLAS and cuFFT. Additionally, our FFT kernel integrates a built-in high-frequency truncation, input zero-padding, and pruning feature to avoid additional memory copy kernels. To fuse the FFT and GEMM workloads, we propose an FFT variant in which a single thread block iterates over the hidden dimension, aligning with the k-loop in GEMM. Additionally, we design two shared memory swizzling patterns to achieve 100% memory bank utilization when forwarding FFT output to GEMM and enabling the iFFT to retrieve GEMM results directly from shared memory. Experimental result on an NVIDIA A100 GPU shows TurboFNO outperforms PyTorch, cuBLAS, and cuFFT by up to 150%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor CoresDaniel Y. Fu, Hermann Kumbong, Eric Nguyen, Christopher RéICLR 2024 · 被引用 41 次
- High-performance Effective Scientific Error-bounded Lossy Compression with Auto-tuned Multi-component InterpolationJinyang Liu, Sheng Di, Kai Zhao, Xin Liang 等SIGMOD 2024 · 被引用 29 次
- cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationJinyang Liu, Jiannan Tian, Shixun Wu, Sheng Di 等SC 2024 · 被引用 17 次
相关 Paper
- Guaranteed Approximation Bounds for Mixed-Precision Neural OperatorsRenbo Tu, Colin White, Jean Kossaifi, Boris Bonev 等ICLR 2024 · 被引用 10 次
- Exploring Efficient Partial Differential Equation Solution Using Speed Galerkin TransformerXun Wang, Zeyang Zhu, Xiangyu Meng, Tao SongSC 2024
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 被引用 27 次
- TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUsShixun Wu, Yujia Zhai, Jinyang Liu, Jiajun Huang 等PPoPP 2025 · 被引用 7 次
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong 等ISCA 2026
