HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration
Guang Fan, Yi Chen, Lei Chen, Liang Kong, Chao Niu, Dian Jiao, Yilan Zhu, Geng Yang, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng
Abstract
The escalating demand for encrypted data computation emphasizes the importance of Fully Homomorphic Encryption (FHE). With a massively parallel architecture, GPUs have emerged as de facto accelerators for practical FHE deployment. However, the performance of existing GPU-based FHE solutions remains severely limited by frequent memory accesses, leading to under-utilized computing capability. To identify the performance bottlenecks, we conduct detailed profiling experiments across both the Number Theoretic Transform (NTT) arithmetic level and the bootstrapping level. At the arithmetic level, we observe that GPU instruction pipeline stalls are incurred by various memory access patterns within the inner and outer parts of a divided NTT. Meanwhile, redundant data movements between on-chip memory and off-chip global memory are exhibited in the conventional step-by-step GPU kernel execution mode of FHE operations. Based on these observations, we introduce HyperDrive, a hierarchical optimization strategy towards arithmetic and operation level, respectively, to exploit memory access efficiency in GPU-based FHE. Specifically, we first design a fine-grained register data access method in each FP64 Tensor Core thread for the inner part of NTT to reduce frequent interactions with shared memory. Secondly, we develop a structured coalescing scheme for off-chip global memory access for the outer part of NTT. Last but not least, at the operation level, we propose a memory-aware co-optimization for NTT and the cross-poly kernels. To further mitigate bottlenecks, we employ a ciphertextreusing multiply-and-accumulate within plaintext and ciphertext. Compared with representative solutions, WarpDrive and Neo, HyperDrive delivers speedups of to in NTT, to 3.7× in homomorphic multiplication, and 3.7× to 8.2× in endto-end workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cf518ff3-93c5-4a44-af92-e72f5061f452Related papers
- WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresGuang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan et al.HPCA 2025 · 29 citations
- MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access OptimizationJunyi Zhang, Xianglong Deng, Yi Chen, Guang Fan et al.ISCA 2026
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou et al.HPCA 2023 · 90 citations
- An NTT/INTT Accelerator with Ultra-High Throughput and Area Efficiency for FHEZhaojun Lu, Weizong Yu, Peng Xu, Wei Wang et al.DAC 2024 · 3 citations
