Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted Inference
Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan, Wenting Zheng, Dimitrios Skarlatos
Abstract
Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies that limit their accessibility. GPUs offer a more widely available alternative, yet achieving ASIC-class performance on GPUs is challenging. Large models such as LLMs compound these challenges by requiring optimized kernels, terabyte-scale memory management, and efficient execution across multiple devices.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f321fb93-cd40-4cc7-936d-f7fc85e693f1Related papers
- EncryptedLLM: Privacy-Preserving Large Language Model Inference via GPU-Accelerated Fully Homomorphic EncryptionLeo de Castro, Daniel Escudero, Adya Agrawal, Antigoni Polychroniadou et al.ICML 2025
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi et al.HPCA 2025 · 14 citations
- Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption AcceleratorLiang Kong, Xianglong Deng, Guang Fan, Shengyu Fan et al.ASPLOS 2026
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen et al.MICRO 2023 · 46 citations
- BitPacker: Enabling High Arithmetic Efficiency in Fully Homomorphic Encryption AcceleratorsNikola Samardzic, Daniel SánchezASPLOS 2024 · 19 citations
