Guardain: Protecting Emerging Generative AI Workloads on Heterogeneous NPU
Aritra Dhar, Clément Thorens, Lara Magdalena Lazier, Lukas Cavigelli
Abstract
Driven by recent advances in large language models (LLMs), generative AI applications have become the dominant workload for the modern cloud. Specialized hardware accelerators, such as GPUs, NPUs, and TPUs, play a key role in AI adoption due to their superior performance over general-purpose CPUs. AI models and the data are often highly sensitive and come from mutually distrusting parties. Existing industry-standard CPU-based TEEs, such as Intel SGX or AMD SEV, do not adequately protect these accelerators. Device-TEEs like Nvidia-CC only address tightly coupled CPU-GPU systems with a proprietary solution requiring TEE on the host CPU side. On the other hand, existing academic proposals target specific CPU-TEE platforms. To address this gap, we propose Guardain,a confidential computing architecture for discrete NPU devices that requires no trust in the host system. Guardainsecures data, model parameters, and operator binaries through authenticated encryption. Guardainuses delegation-based memory semantics to ensure isolation from the host software stack, while task attestation guarantees strong model integrity. Our G Uardainimplementation and evaluation with state-of-the-art LLMs such as Llama2 and Llama3 shows that Guardainintroduces minimal overhead with no changes in the AI software stack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b280e6cb-4c73-43fc-8eba-48d39c24fc69Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order ExecutionJo Van Bulck, Marina Minkin, Ofir Weisse, Daniel Genkin et al.USENIX Security 2018 · 1,175 citations
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- CrypTFlow: Secure TensorFlow InferenceNishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta et al.S&P 2020 · 276 citations
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPUSijun Tan, Brian Knott, Yuan Tian, David J. WuS&P 2021 · 241 citations
- Lord of the Ring(s): Side Channel Attacks on the CPU On-Chip Ring Interconnect Are PracticalRiccardo Paccagnella, Licheng Luo, Christopher W. FletcherUSENIX Security 2021 · 121 citations
Related papers
- GuardNN: secure accelerator architecture for privacy-preserving deep learningWeizhe Hua, Muhammad Umar, Zhiru Zhang, G. Edward SuhDAC 2022 · 28 citations
- Confidential Computing within an AI AcceleratorKapil Vaswani, Stavros Volos, Cédric Fournet, Antonio Nino Diaz et al.USENIX ATC 2023 · 31 citations
- ACAI: Protecting Accelerator Execution with Arm Confidential Computing ArchitectureSupraja Sridhara, Andrin Bertschi, Benedict Schlüter, Mark Kuhne et al.USENIX Security 2024 · 36 citations
- Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentJianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang et al.S&P 2020 · 95 citations
- ccAI: A Compatible and Confidential System for AI ComputingChenxu Wang, Danqing Tang, Changxu Ci, Junjie Huang et al.MICRO 2025 · 3 citations
