TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU
Shengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou, Dan Meng, Mingzhe Zhang
摘要
In the cloud computing era, privacy protection is becoming pervasive in a broad range of applications (e.g., machine learning, data mining, etc). Fully Homomorphic Encryption (FHE) is considered the perfect solution as it enables privacypreserved computation on untrusted servers. Unfortunately, the prohibitive performance overhead blocks the wide adoption of FHE (about 10, 000× slower than the normal computation). As heterogeneous architectures have gained remarkable success in several fields, achieving high performance for FHE with specifically designed accelerators seems to be a natural choice. Until now, most FHE accelerators have focused on efficiently implementing one FHE operation at a time based on ASIC and with significantly higher performance than GPU and FPGA. However, recent state-of-the-art FHE accelerators rely on an expensive and large on-chip storage and a high-end manufacturing process (i.e., 7nm), which increase the cost of FHE adoption.
In this paper, we propose TensorFHE, an FHE acceleration solution based on GPGPU for real applications on encrypted data. TensorFHE utilizes Tensor Core Units (TCUs) to boost the computation of Number Theoretic Transform (NTT), which is the part of FHE with highest time-cost. Moreover, TensorFHE focuses on performing as many FHE operations as possible in a certain time period rather than reducing the latency of one operation. Based on such an idea, TensorFHE introduces operation-level batching to fully utilize the data parallelism in GPGPU. We experimentally prove that it is possible to achieve comparable performance with GPGPU as with state-of-the-art ASIC accelerators. TensorFHE performs 913 KOPS and 88 KOPS for NTT and HMULT (key FHE kernels) within NVIDIA A100 GPGPU, which is 2.61× faster than state-of-the-art FHE implementation on GPGPU; Moreover, TensorFHE provides comparable performance to the ASIC FHE accelerators, which makes it even 2.9× faster than the F1+ with a specific workload. Such a pure software acceleration based on commercial hardware with high performance can open up usage of state-of-the-art FHE algorithms for a broad set of applications in real systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen 等MICRO 2023 · 被引用 46 次
- Trinity: A General Purpose FHE AcceleratorXianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian 等MICRO 2024 · 被引用 34 次
- PrivCirNet: Efficient Private Inference via Block Circulant TransformationTianshi Xu, Lemeng Wu, Runsheng Wang, Meng LiNeurIPS 2024 · 被引用 21 次
- GPU-based Private Information Retrieval for On-Device Machine Learning InferenceMaximilian Lam, Jeff Johnson, Wenjie Xiong, Kiwan Maeng 等ASPLOS 2024 · 被引用 11 次
- Cinnamon: A Framework for Scale-Out Encrypted AISiddharth Jayashankar, Edward Chen, Tom Tang, Wenting Zheng 等ASPLOS 2025 · 被引用 10 次
它引用的顶会 Paper9
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas 等MICRO 2021 · 被引用 294 次
- HEAX: An Architecture for Computing on Encrypted DataM. Sadegh Riazi, Kim Laine, Blake Pelton, Wei DaiASPLOS 2020 · 被引用 244 次
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar 等ISCA 2022 · 被引用 205 次
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung 等ISCA 2022 · 被引用 184 次
相关 Paper
- WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresGuang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan 等HPCA 2025 · 被引用 29 次
- Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor CoreDian Jiao, Xianglong Deng, Zhiwei Wang, Shengyu Fan 等ISCA 2025 · 被引用 18 次
- MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access OptimizationJunyi Zhang, Xianglong Deng, Yi Chen, Guang Fan 等ISCA 2026
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi 等HPCA 2025 · 被引用 14 次
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong 等ISCA 2026
