Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted Inference
Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan, Wenting Zheng, Dimitrios Skarlatos
摘要
Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies that limit their accessibility. GPUs offer a more widely available alternative, yet achieving ASIC-class performance on GPUs is challenging. Large models such as LLMs compound these challenges by requiring optimized kernels, terabyte-scale memory management, and efficient execution across multiple devices.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- EncryptedLLM: Privacy-Preserving Large Language Model Inference via GPU-Accelerated Fully Homomorphic EncryptionLeo de Castro, Daniel Escudero, Adya Agrawal, Antigoni Polychroniadou 等ICML 2025
- Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in MemoryJongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi 等HPCA 2025 · 被引用 14 次
- Falcon: Algorithm-Hardware Co-Design for Efficient Fully Homomorphic Encryption AcceleratorLiang Kong, Xianglong Deng, Guang Fan, Shengyu Fan 等ASPLOS 2026
- GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic EncryptionKaustubh Shivdikar, Yuhui Bao, Rashmi Agrawal, Michael Tian Shen 等MICRO 2023 · 被引用 46 次
- BitPacker: Enabling High Arithmetic Efficiency in Fully Homomorphic Encryption AcceleratorsNikola Samardzic, Daniel SánchezASPLOS 2024 · 被引用 19 次
