LÆGIS: Pinpointing and Addressing Performance Overheads of GPU-Based Confidential Computing
Yang Yang, Adwait Jog
Abstract
GPU-based Confidential Computing (CC) combines a CPU trusted execution environment (TEE) with a CC-capable GPU to protect data in use. However, it faces a memory wall: GPU memory is far more limited than CPU memory. Unified virtual memory (UVM) alleviates this constraint by allowing the CPU and GPU to share a single address space, with pages managed jointly by the GPU's memory management unit (GMMU) and a CPU-side UVM driver. This design, however, is fault-driven and introduces significant overhead - when the GPU accesses pages residing in CPU memory, it triggers page faults that are aggregated by the GMMU and processed in batches by the driver. Under CC, every page migrated in this process must additionally be encrypted to ensure confidentiality and integrity, further amplifying the cost. We identify three key inefficiencies in UVM-based GPU CC. First, it requires tight CPU-GPU synchronization to negotiate initialization vectors (IVs) for AES-GCM encryption, placing encryption on the critical path. While this design avoids storing IVs and thus eliminates the overhead of integrity-tree maintenance, the resulting synchronization overhead significantly degrades performance. Second, the driver thread often sits idle while waiting for new fault batches to arrive, wasting CPU cycles that could otherwise be used for encryption. Third, even when the CPU is utilized, CPU software encryption throughput is low (1.3 GB/s) because UVM relies on Linux Kernel Crypto APIs that are not currently parallelized, further compounding the slowdown. To address these inefficiencies, we propose LÆGIS, a design that opportunistically pre-encrypts pages on the CPU. LÆGIS leverages secure high-bandwidth memory (HBM) for flexible IV management, introducing an IV Bank stored in 3D-stacked HBM that decouples encryption from CPU-GPU synchronization while eliminating the need for integrity trees. This enables out-of-order encryption and substantially improves UVM performance under CC. Our evaluation shows that LÆGIS significantly reduces CC overhead, achieving up to on average) and (2.74× on average) speedup over the CC baseline under default and aggressive prefetching, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 9 citations
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 73 citations
- SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine LearningJoongun Park, Yongqin Wang, Huan Xu, Hanjiang Wu et al.HPCA 2026
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 7 citations
- CAGE: Complementing Arm CCA with GPU ExtensionsChenxu Wang, Fengwei Zhang, Yunjie Deng, Kevin Leach et al.NDSS 2024
