HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan, Yi Chen, Lei Chen, Liang Kong, Chao Niu, Dian Jiao, Yilan Zhu, Geng Yang, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng
摘要
The escalating demand for encrypted data computation emphasizes the importance of Fully Homomorphic Encryption (FHE). With a massively parallel architecture, GPUs have emerged as de facto accelerators for practical FHE deployment. However, the performance of existing GPU-based FHE solutions remains severely limited by frequent memory accesses, leading to under-utilized computing capability. To identify the performance bottlenecks, we conduct detailed profiling experiments across both the Number Theoretic Transform (NTT) arithmetic level and the bootstrapping level. At the arithmetic level, we observe that GPU instruction pipeline stalls are incurred by various memory access patterns within the inner and outer parts of a divided NTT. Meanwhile, redundant data movements between on-chip memory and off-chip global memory are exhibited in the conventional step-by-step GPU kernel execution mode of FHE operations. Based on these observations, we introduce HyperDrive, a hierarchical optimization strategy towards arithmetic and operation level, respectively, to exploit memory access efficiency in GPU-based FHE. Specifically, we first design a fine-grained register data access method in each FP64 Tensor Core thread for the inner part of NTT to reduce frequent interactions with shared memory. Secondly, we develop a structured coalescing scheme for off-chip global memory access for the outer part of NTT. Last but not least, at the operation level, we propose a memory-aware co-optimization for NTT and the cross-poly kernels. To further mitigate bottlenecks, we employ a ciphertextreusing multiply-and-accumulate within plaintext and ciphertext. Compared with representative solutions, WarpDrive and Neo, HyperDrive delivers speedups of to in NTT, to 3.7× in homomorphic multiplication, and 3.7× to 8.2× in endto-end workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresGuang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan 等HPCA 2025 · 被引用 29 次
- MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access OptimizationJunyi Zhang, Xianglong Deng, Yi Chen, Guang Fan 等ISCA 2026
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi 等HPCA 2025 · 被引用 14 次
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou 等HPCA 2023 · 被引用 90 次
- An NTT/INTT Accelerator with Ultra-High Throughput and Area Efficiency for FHEZhaojun Lu, Weizong Yu, Peng Xu, Wei Wang 等DAC 2024 · 被引用 3 次
