Supporting Secure Multi-GPU Computing with Dynamic and Batched Metadata Management
Seonjin Na, Jungwoo Kim, Sunho Lee, Jaehyuk Huh
摘要
With growing problem sizes for GPU computing, multi-GPU systems with fine-grained memory sharing have emerged to improve the current coarse-grained unified memory support based on page migration. Such multi-GPU systems with shared memory pose a new challenge in securing CPU-GPU and inter-GPU communications, as the cost of secure data transfers adds a significant performance overhead. There are two overheads of secure communication in multi-GPU systems: First, extra overhead is added to generate one-time pads (OTPs) for authenticated encryption. Second, the security metadata such as MACs and counters passed along with encrypted data consume precious network bandwidth. This study investigates the performance impact of secure communication in multi-GPU systems and evaluates the prior CPU-oriented OTP precomputation schemes adapted for multi-GPU systems. Our investigation identifies the challenge with the limited OTP buffers for inter-GPU communication and the opportunity to reduce traffic for security meta-data with bursty communications in GPUs. Based on the analysis, this paper proposes a new dynamic OTP buffer allocation technique, which adjusts the buffer assignment for each source-destination pair to reflect the communication patterns. To address the bandwidth problem by extra security metadata, the study employs a dynamic batching scheme to transfer only a single set of metadata for each batched group of data responses. The proposed design constantly tracks the communication pattern from each GPU, periodically adjusts the allocated buffer size, and dynamically forms batches of data transfers. Our evaluation shows that in a 16-GPU system, the proposed scheme can improve the performance by 13.2 % and 17.5 % on average from the prior cached and private schemes, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient Security Support for CXL Memory through Adaptive Incremental Offloaded (Re-)EncryptionChuanhan Li, Jishen Zhao, Yuanchao XuMICRO 2025 · 被引用 5 次
- Unified Memory Protection with Multi-granular MAC and Integrity Tree for Heterogeneous ProcessorsSunho Lee, Seonjin Na, Jeongwon Choi, Jinwon Pyo 等ISCA 2025 · 被引用 2 次
- SoK: Analysis of Accelerator TEE DesignsChenxu Wang, Junjie Huang, Yujun Liang, Xuanyao Peng 等NDSS 2026 · 被引用 2 次
它引用的顶会 Paper16
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin 等S&P 2019 · 被引用 2,435 次
- Meltdown: Reading Kernel Memory from User SpaceMoritz Lipp, Michael Schwarz, Daniel Gruss, Thomas Prescher 等USENIX Security 2018 · 被引用 1,456 次
- Scalable Memory Protection in the PENGLAI EnclaveErhu Feng, Xu Lu, Dong Du, Bicheng Yang 等OSDI 2021 · 被引用 126 次
- Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentJianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang 等S&P 2020 · 被引用 95 次
- Osiris: Automated Discovery of Microarchitectural Side ChannelsDaniel Weber, Ahmad Ibrahim, Hamed Nemati, Michael Schwarz 等USENIX Security 2021 · 被引用 75 次
相关 Paper
- Plutus: Bandwidth-Efficient Memory Security for GPUsRahaf Abdullah, Huiyang Zhou, Amro AwadHPCA 2023 · 被引用 13 次
- Common Counters: Compressed Encryption Counters for Secure GPU MemorySeonjin Na, Sunho Lee, Yeonjae Kim, Jongse Park 等HPCA 2021 · 被引用 34 次
- Salus: Efficient Security Support for CXL-Expanded GPU MemoryRahaf Abdullah, Hyokeun Lee, Huiyang Zhou, Amro AwadHPCA 2024 · 被引用 16 次
- Adaptive Security Support for Heterogeneous Memory on GPUsShougang Yuan, Amro Awad, Ardhi Wiratama Baskara Yudha, Yan Solihin 等HPCA 2022 · 被引用 15 次
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang 等HPCA 2024 · 被引用 19 次
