Lune

ISCA2026Top-tier venue

HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE Acceleration

Guang Fan, Yi Chen, Lei Chen, Liang Kong, Chao Niu, Dian Jiao, Yilan Zhu, Geng Yang, Shengyu Fan, Xianglong Deng, Fangyu Zheng, Jian Weng

2026Year

Abstract

The escalating demand for encrypted data computation emphasizes the importance of Fully Homomorphic Encryption (FHE). With a massively parallel architecture, GPUs have emerged as de facto accelerators for practical FHE deployment. However, the performance of existing GPU-based FHE solutions remains severely limited by frequent memory accesses, leading to under-utilized computing capability. To identify the performance bottlenecks, we conduct detailed profiling experiments across both the Number Theoretic Transform (NTT) arithmetic level and the bootstrapping level. At the arithmetic level, we observe that GPU instruction pipeline stalls are incurred by various memory access patterns within the inner and outer parts of a divided NTT. Meanwhile, redundant data movements between on-chip memory and off-chip global memory are exhibited in the conventional step-by-step GPU kernel execution mode of FHE operations. Based on these observations, we introduce HyperDrive, a hierarchical optimization strategy towards arithmetic and operation level, respectively, to exploit memory access efficiency in GPU-based FHE. Specifically, we first design a fine-grained register data access method in each FP64 Tensor Core thread for the inner part of NTT to reduce frequent interactions with shared memory. Secondly, we develop a structured coalescing scheme for off-chip global memory access for the outer part of NTT. Last but not least, at the operation level, we propose a memory-aware co-optimization for NTT and the cross-poly kernels. To further mitigate bottlenecks, we employ a ciphertextreusing multiply-and-accumulate within plaintext and ciphertext. Compared with representative solutions, WarpDrive and Neo, HyperDrive delivers speedups of 2.0×\mathbf{2. 0} \times to 5.5×\mathbf{5. 5} \times in NTT, 2.2×\mathbf{2. 2} \times to 3.7× in homomorphic multiplication, and 3.7× to 8.2× in endto-end workloads.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get cf518ff3-93c5-4a44-af92-e72f5061f452

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines