WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA Cores
Guang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan, Tian Zhou, Xianglong Deng, Wenxu Tang, Liang Kong, Yixuan Song, Shoumeng Yan
Abstract
The application of Fully Homomorphic Encryption (FHE) is rapidly gaining traction as a means to maintain data confidentiality while performing computations on encrypted data. Given the accessibility and computational power, GPUs hold promise for significantly accelerating FHE operations. However, existing GPU-based acceleration solutions face several formidable challenges, notably the extensive occurrence of pipeline stalls induced by memory access and suboptimal harnessing of GPU hardware. This paper presents WarpDrive, a comprehensive framework for GPU-based FHE acceleration. Through sophisticated computation decomposition and fine-grained memory access design, WarpDrive significantly reduces the number of instructions by and pipeline stalls by compared to the state-of-the-art solution. Additionally, WarpDrive features a framework that supports the concurrent utilization of CUDA Cores and Tensor Cores within the NTT operation, for the first time, achieving performance that surpasses that of any single type of processing unit. Furthermore, we fully exploit the intra-ciphertext parallelism to elevate both computation and memory utilization, achieving up to improvements without the need for ciphertext batching. Experimental results demonstrate that our optimizations highly enhance the performance of homomorphic operations. On an NVIDIA A100 GPU, WarpDrive achieves a throughput of 1218 KOPS for NTT and 305 KOPS for homomorphic multiplication, outperforming the state-of-the-art GPU solution (TensorFHE) by factors of and , respectively. For the specific FHE workload, even under a much smaller batch size, our approach achieves the performance of TensorFHE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get be61d129-d0fb-4afa-96f7-23cb14515939Cited by top-tier papers4
- Cheddar: A Swift Fully Homomorphic Encryption Library Designed for GPU ArchitecturesWonseok Choi, Jongmin Kim, Jung Ho AhnASPLOS 2026 · 6 citations
- Leveraging ASIC AI Chips for Homomorphic EncryptionJianming Tong, Tianhao Huang, Jingtian Dang, Leo de Castro et al.HPCA 2026 · 2 citations
- Fenc2: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment EncodingRan Ran, Zhaoting Gong, Nuo Xu, Yuanchao Xu et al.ISCA 2026
- BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV CacheDayou Du, Shijie Cao, Jianyi Cheng, Luo Mai et al.HPCA 2026
Related papers
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong et al.ISCA 2026
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou et al.HPCA 2023 · 90 citations
- MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access OptimizationJunyi Zhang, Xianglong Deng, Yi Chen, Guang Fan et al.ISCA 2026
- Strix: An End-to-End Streaming Architecture with Two-Level Ciphertext Batching for Fully Homomorphic Encryption with Programmable BootstrappingAdiwena Putra, Prasetiyo, Yi Chen, John Kim et al.MICRO 2023 · 27 citations
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
