TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU
Shengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou, Dan Meng, Mingzhe Zhang
Abstract
In the cloud computing era, privacy protection is becoming pervasive in a broad range of applications (e.g., machine learning, data mining, etc). Fully Homomorphic Encryption (FHE) is considered the perfect solution as it enables privacypreserved computation on untrusted servers. Unfortunately, the prohibitive performance overhead blocks the wide adoption of FHE (about 10, 000× slower than the normal computation). As heterogeneous architectures have gained remarkable success in several fields, achieving high performance for FHE with specifically designed accelerators seems to be a natural choice. Until now, most FHE accelerators have focused on efficiently implementing one FHE operation at a time based on ASIC and with significantly higher performance than GPU and FPGA. However, recent state-of-the-art FHE accelerators rely on an expensive and large on-chip storage and a high-end manufacturing process (i.e., 7nm), which increase the cost of FHE adoption.
In this paper, we propose TensorFHE, an FHE acceleration solution based on GPGPU for real applications on encrypted data. TensorFHE utilizes Tensor Core Units (TCUs) to boost the computation of Number Theoretic Transform (NTT), which is the part of FHE with highest time-cost. Moreover, TensorFHE focuses on performing as many FHE operations as possible in a certain time period rather than reducing the latency of one operation. Based on such an idea, TensorFHE introduces operation-level batching to fully utilize the data parallelism in GPGPU. We experimentally prove that it is possible to achieve comparable performance with GPGPU as with state-of-the-art ASIC accelerators. TensorFHE performs 913 KOPS and 88 KOPS for NTT and HMULT (key FHE kernels) within NVIDIA A100 GPGPU, which is 2.61× faster than state-of-the-art FHE implementation on GPGPU; Moreover, TensorFHE provides comparable performance to the ASIC FHE accelerators, which makes it even 2.9× faster than the F1+ with a specific workload. Such a pure software acceleration based on commercial hardware with high performance can open up usage of state-of-the-art FHE algorithms for a broad set of applications in real systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89341cde-ba58-4168-9f5b-dfa57167a66aCited by top-tier papers12
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen et al.MICRO 2023 · 46 citations
- Trinity: A General Purpose FHE AcceleratorXianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian et al.MICRO 2024 · 34 citations
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 21 citations
- GPU-based Private Information Retrieval for On-Device Machine Learning InferenceMaximilian Lam, Jeff Johnson, Wenjie Xiong, Kiwan Maeng et al.ASPLOS 2024 · 11 citations
- Cinnamon: A Framework for Scale-Out Encrypted AISiddharth Jayashankar, Edward Chen, Tom Tang, Wenting Zheng et al.ASPLOS 2025 · 10 citations
Builds on9
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas et al.MICRO 2021 · 294 citations
- HEAX: An Architecture for Computing on Encrypted DataM. Sadegh Riazi, Kim Laine, Blake Pelton, Wei DaiASPLOS 2020 · 244 citations
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar et al.ISCA 2022 · 205 citations
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung et al.ISCA 2022 · 184 citations
Related papers
- WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresGuang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan et al.HPCA 2025 · 29 citations
- Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor CoreDian Jiao, Xianglong Deng, Zhiwei Wang, Shengyu Fan et al.ISCA 2025 · 18 citations
- MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access OptimizationJunyi Zhang, Xianglong Deng, Yi Chen, Guang Fan et al.ISCA 2026
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong et al.ISCA 2026
