GPU Travelling: Efficient Confidential Collaborative Training with TEE-Enabled GPUs
Shixuan Zhao, Zhongshu Gu, Salman Ahmed, Enriquillo Valdez, Hani Jamjoom, Zhiqiang Lin
摘要
Confidential collaborative machine learning (ML) enables multiple mutually distrusted data holders to jointly train an ML model while preserving the confidentiality of their private datasets due to regulatory or competitive reasons. However, existing works need frequent data and model exchanges during training via slower conventional links. They face increasing challenges due to the exponentially growing sizes of models and datasets in modern training workloads like large language models (LLMs), resulting in prohibitively high communication costs. In this paper, we propose a novel mechanism called GPU Travelling that leverages recently emerged confidential GPUs. With our rigorous design, the GPU can securely travel to the specific data holder to load the dataset directly into the GPU's protected memory and then return for training, eliminating the need for data transmission while ensuring confidentiality up to a data-centre level. We developed a prototype using Intel TDX and NVIDIA H100 and evaluated its performance on llm.c, a CUDA-based LLM training project, and demonstrated the performance and feasibility while maintaining strong security guarantees. The results showed at least 4x speed improvement when transmitting a 512 MiB dataset chunk versus conventional transmission.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- vSGX: Virtualizing SGX Enclaves on AMD SEVShixuan Zhao, Mengyuan Li, Yinqian Zhang, Zhiqiang LinS&P 2022 · 被引用 32 次
- CaPC Learning: Confidential and Private Collaborative LearningChristopher A. Choquette-Choo, Natalie Dullerud, Adam Dziedzic, Yunxiang Zhang 等ICLR 2021 · 被引用 24 次
- Understanding Routable PCIe Performance for Composable InfrastructuresWentao Hou, Jie Zhang, Zeke Wang, Ming LiuNSDI 2024 · 被引用 23 次
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang 等NeurIPS 2024 · 被引用 23 次
- DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy AggregationPau-Chen Cheng, Kevin Eykholt, Zhongshu Gu, Hani Jamjoom 等EuroSys 2024 · 被引用 15 次
相关 Paper
- SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine LearningJoongun Park, Yongqin Wang, Huan Xu, Hanjiang Wu 等HPCA 2026
- PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined EncryptionYifan Tan, Cheng Tan, Zeyu Mi, Haibo ChenASPLOS 2025 · 被引用 10 次
- Guardain: Protecting Emerging Generative AI Workloads on Heterogeneous NPUAritra Dhar, Clément Thorens, Lara Magdalena Lazier, Lukas CavigelliS&P 2025
- Enabling Execution Assurance of Federated Learning at Untrusted ParticipantsXiaoli Zhang, Fengting Li, Zeyu Zhang, Qi Li 等INFOCOM 2020 · 被引用 87 次
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPUSijun Tan, Brian Knott, Yuan Tian, David J. WuS&P 2021 · 被引用 241 次
