GPU Travelling: Efficient Confidential Collaborative Training with TEE-Enabled GPUs
Shixuan Zhao, Zhongshu Gu, Salman Ahmed, Enriquillo Valdez, Hani Jamjoom, Zhiqiang Lin
Abstract
Confidential collaborative machine learning (ML) enables multiple mutually distrusted data holders to jointly train an ML model while preserving the confidentiality of their private datasets due to regulatory or competitive reasons. However, existing works need frequent data and model exchanges during training via slower conventional links. They face increasing challenges due to the exponentially growing sizes of models and datasets in modern training workloads like large language models (LLMs), resulting in prohibitively high communication costs. In this paper, we propose a novel mechanism called GPU Travelling that leverages recently emerged confidential GPUs. With our rigorous design, the GPU can securely travel to the specific data holder to load the dataset directly into the GPU's protected memory and then return for training, eliminating the need for data transmission while ensuring confidentiality up to a data-centre level. We developed a prototype using Intel TDX and NVIDIA H100 and evaluated its performance on llm.c, a CUDA-based LLM training project, and demonstrated the performance and feasibility while maintaining strong security guarantees. The results showed at least 4x speed improvement when transmitting a 512 MiB dataset chunk versus conventional transmission.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1ff4581-751e-4e43-9dd9-1cb9d4b44918Cited by top-tier papers1
Ask how each one uses itBuilds on7
- vSGX: Virtualizing SGX Enclaves on AMD SEVShixuan Zhao, Mengyuan Li, Yinqian Zhang, Zhiqiang LinS&P 2022 · 32 citations
- CaPC Learning: Confidential and Private Collaborative LearningChristopher A. Choquette-Choo, Natalie Dullerud, Adam Dziedzic, Yunxiang Zhang et al.ICLR 2021 · 24 citations
- Understanding Routable PCIe Performance for Composable InfrastructuresWentao Hou, Jie Zhang, Zeke Wang, Ming LiuNSDI 2024 · 23 citations
- SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM TrainingJinda Jia, Cong Xie, Hanlin Lu, Daoce Wang et al.NeurIPS 2024 · 23 citations
- DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy AggregationPau-Chen Cheng, Kevin Eykholt, Zhongshu Gu, Hani Jamjoom et al.EuroSys 2024 · 15 citations
Related papers
- SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine LearningJoongun Park, Yongqin Wang, Huan Xu, Hanjiang Wu et al.HPCA 2026
- PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined EncryptionYifan Tan, Cheng Tan, Zeyu Mi, Haibo ChenASPLOS 2025 · 10 citations
- Guardain: Protecting Emerging Generative AI Workloads on Heterogeneous NPUAritra Dhar, Clément Thorens, Lara Magdalena Lazier, Lukas CavigelliS&P 2025
- Enabling Execution Assurance of Federated Learning at Untrusted ParticipantsXiaoli Zhang, Fengting Li, Zeyu Zhang, Qi Li et al.INFOCOM 2020 · 87 citations
- CryptGPU: Fast Privacy-Preserving Machine Learning on the GPUSijun Tan, Brian Knott, Yuan Tian, David J. WuS&P 2021 · 241 citations
