GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic Encryption
Kaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen, Gilbert Jonatan, Evelio Mora, Alexander Ingare, Neal Livesay, José L. Abellán, John Kim, Ajay Joshi, David R. Kaeli
摘要
Fully Homomorphic Encryption (FHE) enables the processing of encrypted data without decrypting it. FHE has garnered significant attention over the past decade as it supports secure outsourcing of data processing to remote cloud services. Despite its promise of strong data privacy and security guarantees, FHE introduces a slowdown of up to five orders of magnitude as compared to the same computation using plaintext data. This overhead is presently a major barrier to the commercial adoption of FHE.
In this work, we leverage GPUs to accelerate FHE, capitalizing on a well-established GPU ecosystem available in the cloud. We propose GME, which combines three key microarchitectural extensions along with a compile-time optimization to the current AMD CDNA GPU architecture. First, GME integrates a lightweight on-chip compute unit (CU)-side hierarchical interconnect to retain ciphertext in cache across FHE kernels, thus eliminating redundant memory transactions. Second, to tackle compute bottlenecks, GME introduces special MOD-units that provide native custom hardware support for modular reduction operations, one of the most commonly executed sets of operations in FHE. Third, by integrating the MOD-unit with our novel pipelined 64-bit integer arithmetic cores (WMAC-units), GME further accelerates FHE workloads by 19%. Finally, we propose a Locality-Aware Block Scheduler (LABS) that exploits the temporal locality available in FHE primitive blocks. Incorporating these microarchitectural features and compiler optimizations, we create a synergistic approach achieving average speedups of 796×, 14.2×, and 2.3× over Intel Xeon CPU, NVIDIA V100 GPU, and Xilinx FPGA implementations, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial AcceleratorKaustubh Shivdikar, Nicolas Bohm Agostini, Malith Jayaweera, Gilbert Jonatan 等ISCA 2024 · 被引用 9 次
- Cheddar: A Swift Fully Homomorphic Encryption Library Designed for GPU ArchitecturesWonseok Choi, Jongmin Kim, Jung Ho AhnASPLOS 2026 · 被引用 6 次
- Towards Closing the Performance Gap for Cryptographic Kernels Between CPUs and Specialized HardwareNaifeng Zhang, Sophia Fu, Franz FranchettiMICRO 2025 · 被引用 4 次
- Leveraging ASIC AI Chips for Homomorphic EncryptionJianming Tong, Tianhao Huang, Jingtian Dang, Leo de Castro 等HPCA 2026 · 被引用 2 次
- SLOTHE : Lazy Approximation of Non-Arithmetic Neural Network Functions over Encrypted DataKevin Nam, Youyeon Joo, Seungjin Ha, Yunheung PaekUSENIX Security 2025
它引用的顶会 Paper10
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas 等MICRO 2021 · 被引用 294 次
- HEAX: An Architecture for Computing on Encrypted DataM. Sadegh Riazi, Kim Laine, Blake Pelton, Wei DaiASPLOS 2020 · 被引用 244 次
- CraterLake: a hardware accelerator for efficient unbounded computation on encrypted dataNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Nathan Manohar 等ISCA 2022 · 被引用 205 次
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung 等ISCA 2022 · 被引用 184 次
- Efficient Bootstrapping for Approximate Homomorphic Encryption with Non-sparse KeysJean-Philippe Bossuat, Christian Mouchet, Juan Ramón Troncoso-Pastoriza, Jean-Pierre HubauxEUROCRYPT 2021 · 被引用 179 次
相关 Paper
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi 等HPCA 2025 · 被引用 14 次
- Neo: Towards Efficient Fully Homomorphic Encryption Acceleration using Tensor CoreDian Jiao, Xianglong Deng, Zhiwei Wang, Shengyu Fan 等ISCA 2025 · 被引用 18 次
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou 等HPCA 2023 · 被引用 90 次
- Peregrine: Accelerating TFHE Bootstrapping on GPUs via Multi-Level External Product Co-DesignHaoqi He, Zhiwei Wang, Lutan Zhao, Dian Jiao 等HPCA 2026 · 被引用 1 次
- SHARP: A Short-Word Hierarchical Accelerator for Robust and Practical Fully Homomorphic EncryptionJongmin Kim, Sangpyo Kim, Jaewan Choi, Jaiyoung Park 等ISCA 2023 · 被引用 110 次
