MNEMOS: A GPU-Based TFHE Acceleration Framework with Memory Access Optimization
Junyi Zhang, Xianglong Deng, Yi Chen, Guang Fan, Lei Chen, Dian Jiao, Shengyu Fan, Zhiwei Wang, Mingzhe Zhang
Abstract
Fully Homomorphic Encryption over Torus (TFHE) provides a promising approach for privacy-preserving computing by enabling computation directly on encrypted data. However, this strong security guarantee comes at the cost of enormous computational overhead compared with plaintext computation. Modern Graphics Processing Units (GPUs), equipped with thousands of parallel computing cores and high memory bandwidth, offer an attractive platform for accelerating TFHE workloads. By exploiting their massive parallelism, the latency of TFHE primitives can be significantly reduced, making privacy-preserving computing practical. Nevertheless, executing the TFHE applications on the GPU remains limited. In this paper, we propose MNEMOS, a TFHE acceleration framework for GPU platform, optimizing the memory access during TFHE execution. In our study, we observe that severe pipeline stalls occur during TFHE kernel execution, primarily caused by frequent memory accesses and cache misses, which significantly degrade overall performance. Moreover, the utilization of Tensor Cores (TCUs) is far from optimal. Due to the excessive memory access latency caused by frequent memory accesses, the computational throughput of TCUs cannot be fully exploited. To address these issues, we propose a memory-aware algorithmic optimization that improves reuse efficiency through data re-layout and access scheduling. In addition, we introduce a Tensor-Core-optimized FFT mapping strategy that mitigates performance degradation caused by cache misses and enhances the effective utilization of TCUs during PBS computation. Experiments demonstrate that our optimizations highly enhance the performance of TFHE-based applications by on average and up to in the best case.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5a17d3f1-e005-4763-a4b7-1951b56ba11cRelated papers
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou et al.HPCA 2023 · 90 citations
- WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresGuang Fan, Mingzhe Zhang, Fangyu Zheng, Shengyu Fan et al.HPCA 2025 · 29 citations
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
- HyperDrive: Hierarchical Exploitation of Memory Efficiency for GPU-Based FHE AccelerationGuang Fan, Yi Chen, Lei Chen, Liang Kong et al.ISCA 2026
- Affinity-based Optimizations for TFHE on Processing-in-DRAMKevin Nam, Heon Hui Jung, Hyunyoung Oh, Yunheung PaekASPLOS 2025 · 2 citations
