USENIX Security2026Top-tier venue
Behind Bars: A Side-Channel Attack on NVIDIA MIG Cache Partitioning Using Memory Barriers
Cheng Gu, Reese Levine, Zhenkai Zhang, Tyler Sorensen, Yanan Guo
Abstract
NVIDIA Multi-Instance GPU (MIG) is a feature designed to enable isolation and secure multi-tenancy on large data center GPUs. MIG partitions a single GPU into multiple instances, each with dedicated hardware resources such as L2 cache slices. MIG is also documented to form the foundation of NVIDIA's confidential computing stack by providing hardware-isolated trusted execution environments. However, the security claims of MIG deserve closer investigation, especially given the complexity of the GPU memory system and its many (sparsely documented) memory instructions. In this work, we empirically examine the behavior of GPU L2 cache with MIG enabled. We find that despite the partitioning design, cross-instance L2 cache interference still occurs. Specifically, memory barriers (membars) generated in one MIG instance have side effects that propagate across L2 partitions and affect the timing of certain load operations in other instances. We also find that these membars can be triggered by specific GPU activities, such as kernel launches. Building on these observations, we develop a new timing-based side-channel attack in which an attacker in one MIG instance can infer the kernel launch patterns of a victim in another instance. We show that this attack compromises the confidentiality of widely used GPU applications, such as large language model inference, because kernel launch patterns in these applications are correlated with sensitive information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- Rendered Insecure: GPU Side Channel Attacks are PracticalHoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-GhazalehCCS 2018 · 214 citations
- Prime+Abort: A Timer-Free High-Precision L3 Cache Attack using Intel TSXCraig Disselkoen, David Kohlbrenner, Leo Porter, Dean M. TullsenUSENIX Security 2017 · 186 citations
- Robust Website Fingerprinting Through the Cache Occupancy ChannelAnatoly Shusterman, Lachlan Kang, Yarden Haskal, Yosef Meltser et al.USENIX Security 2019 · 159 citations
Related papers
- TunneLs for Bootlegging: Fully Reverse-Engineering GPU TLBs for Challenging Isolation Guarantees of NVIDIA MIGZhenkai Zhang, Tyler N. Allen, Fan Yao, Xing Gao et al.CCS 2023 · 18 citations
- Veiled Pathways: Investigating Covert and Side Channels Within GPU UncoreYuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang et al.MICRO 2024 · 8 citations
- Uncovering Real GPU NoC Characteristics: Implications on Interconnect ArchitectureZhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan et al.MICRO 2024 · 15 citations
- STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUBingyao Li, Yueqi Wang, Tianyu Wang, Lieven Eeckhout et al.MICRO 2024 · 11 citations
- Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU SystemsSankha Baran Dutta, Hoda Naghibijouybari, Arjun Gupta, Nael B. Abu-Ghazaleh et al.ISCA 2023 · 41 citations
