EncryptedLLM: Privacy-Preserving Large Language Model Inference via GPU-Accelerated Fully Homomorphic Encryption
Leo de Castro, Daniel Escudero, Adya Agrawal, Antigoni Polychroniadou, Manuela Veloso
Abstract
As large language models (LLMs) become more powerful, the computation required to run these models is increasingly outsourced to a thirdparty cloud. While this saves clients' computation, it risks leaking the clients' LLM queries to the cloud provider. Fully homomorphic encryption (FHE) presents a natural solution to this problem: simply encrypt the query and evaluate the LLM homomorphically on the cloud machine. The result remains encrypted and can only be learned by the client who holds the secret key. In this work, we propose a GPU-accelerated FHE scheme and leverage it to benchmark an encrypted GPT-2 forward pass. Our approach achieves runtimes that are over 200× faster than the CPU baseline. We also present novel and extensive experimental analysis of approximations of LLM activation functions to maintain accuracy while achieving this performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5469e5aa-6fe8-45ae-8407-e57766a97f02Cited by top-tier papers3
- OSNIP: Balancing the Privacy-Utility-Efficiency Trilemma in LLM Inference via Obfuscated Semantic Null SpaceZhiyuan Cao, Zeyu Ma, Chenhao Yang, HAN ZHENG et al.ICML 2026 · 1 citation
- Hyperion: Private Token Sampling with Homomorphic EncryptionLawrence Lim, Jiaming Liu, Vikas Kalagi, Divyakant Agrawal et al.ACL 2026 · 1 citation
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang et al.USENIX Security 2026
Builds on9
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- Labeled PSI from Fully Homomorphic Encryption with Malicious SecurityHao Chen, Zhicong Huang, Kim Laine, Peter RindalCCS 2018 · 242 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- Efficient Bootstrapping for Approximate Homomorphic Encryption with Non-sparse KeysJean-Philippe Bossuat, Christian Mouchet, Juan Ramón Troncoso-Pastoriza, Jean-Pierre HubauxEUROCRYPT 2021 · 179 citations
- Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionItamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov et al.ICML 2024 · 26 citations
Related papers
- Encryption-Friendly LLM ArchitectureDonghwan Rho, Taeseong Kim, Minje Park, Jung Woo Kim et al.ICLR 2025
- Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted InferenceSiddharth Jayashankar, Joshua Kim, Michael B. Sullivan, Wenting Zheng et al.SOSP 2026
- cuFHEDB: GPU-Accelerated Fully Homomorphic Encryption DatabaseShijie Gao, Feng Zhang, Qian Xu, Yang Li et al.ICDE 2026
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen et al.MICRO 2023 · 46 citations
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
