Behind Bars: A Side-Channel Attack on NVIDIA MIG Cache Partitioning Using Memory Barriers
Cheng Gu, Reese Levine, Zhenkai Zhang, Tyler Sorensen, Yanan Guo
摘要
NVIDIA Multi-Instance GPU (MIG) is a feature designed to enable isolation and secure multi-tenancy on large data center GPUs. MIG partitions a single GPU into multiple instances, each with dedicated hardware resources such as L2 cache slices. MIG is also documented to form the foundation of NVIDIA's confidential computing stack by providing hardware-isolated trusted execution environments. However, the security claims of MIG deserve closer investigation, especially given the complexity of the GPU memory system and its many (sparsely documented) memory instructions. In this work, we empirically examine the behavior of GPU L2 cache with MIG enabled. We find that despite the partitioning design, cross-instance L2 cache interference still occurs. Specifically, memory barriers (membars) generated in one MIG instance have side effects that propagate across L2 partitions and affect the timing of certain load operations in other instances. We also find that these membars can be triggered by specific GPU activities, such as kernel launches. Building on these observations, we develop a new timing-based side-channel attack in which an attacker in one MIG instance can infer the kernel launch patterns of a victim in another instance. We show that this attack compromises the confidentiality of widely used GPU applications, such as large language model inference, because kernel launch patterns in these applications are correlated with sensitive information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah 等ISCA 2024 · 被引用 282 次
- Rendered Insecure: GPU Side Channel Attacks are PracticalHoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-GhazalehCCS 2018 · 被引用 214 次
- Prime+Abort: A Timer-Free High-Precision L3 Cache Attack using Intel TSXCraig Disselkoen, David Kohlbrenner, Leo Porter, Dean M. TullsenUSENIX Security 2017 · 被引用 186 次
- Robust Website Fingerprinting Through the Cache Occupancy ChannelAnatoly Shusterman, Lachlan Kang, Yarden Haskal, Yosef Meltser 等USENIX Security 2019 · 被引用 159 次
相关 Paper
- TunneLs for Bootlegging: Fully Reverse-Engineering GPU TLBs for Challenging Isolation Guarantees of NVIDIA MIGZhenkai Zhang, Tyler N. Allen, Fan Yao, Xing Gao 等CCS 2023 · 被引用 18 次
- Veiled Pathways: Investigating Covert and Side Channels Within GPU UncoreYuanqing Miao, Yingtian Zhang, Dinghao Wu, Danfeng Zhang 等MICRO 2024 · 被引用 8 次
- Uncovering Real GPU NoC Characteristics: Implications on Interconnect ArchitectureZhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan 等MICRO 2024 · 被引用 15 次
- STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUBingyao Li, Yueqi Wang, Tianyu Wang, Lieven Eeckhout 等MICRO 2024 · 被引用 11 次
- Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU SystemsSankha Baran Dutta, Hoda Naghibijouybari, Arjun Gupta, Nael B. Abu-Ghazaleh 等ISCA 2023 · 被引用 41 次
