SCALE: Tackling Communication Bottlenecks in Confidential Distributed Machine Learning
Joongun Park, Yongqin Wang, Huan Xu, Hanjiang Wu, Mengyuan Li, Tushar Krishna
Abstract
Machine Learning (ML) has become a cornerstone in numerous applications, creating the need for secure and efficient distributed ML frameworks. However, maintaining data privacy in these systems poses significant challenges, particularly in distributed environments where user data and model parameters must frequently be transmitted between GPUs. Confidential GPU computing technologies, such as NVIDIA's Confidential Computing (CC) mode, offer hardware-based enterprise solutions designed to protect ML workloads in untrusted environments (e.g., public clouds). These technologies leverage heterogeneous systems that combine Confidential Virtual Machines (CVMs) with GPU-based Trusted Execution Environments (TEEs). Nevertheless, confidential computing introduces considerable performance overhead due to its complex heterogeneous architecture and the high-throughput data flows required across TEE security boundaries. For example, encrypted communication occurs both between CVMs and GPU TEEs, and among multiple GPU TEEs, resulting in significant latency compared to native PCIe or high-speed interconnects such as NVLink. Our extensive evaluation shows that these overheads become particularly severe during collective communication operations, which suffer from encryption-induced delays that negatively impact end-to-end training performance. To address this, we propose a co-encryption design that leverages underutilized GPU resources, optimizes encryption and authentication, and introduces a communication algorithm tailored for confidential settings. We evaluate our design using real ML workloads and execution traces collected from four HGX H100/H200 clusters. While CC mode was not available on current NVIDIA software stacks, we incorporate encryptionaware modeling based on hardware specifications to estimate secure communication overheads. Our results demonstrate areduction in communication-related security costs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- GPU Travelling: Efficient Confidential Collaborative Training with TEE-Enabled GPUsShixuan Zhao, Zhongshu Gu, Salman Ahmed, Enriquillo Valdez et al.CCS 2025
- Common Counters: Compressed Encryption Counters for Secure GPU MemorySeonjin Na, Sunho Lee, Yeonjae Kim, Jongse Park et al.HPCA 2021 · 34 citations
- Guardain: Protecting Emerging Generative AI Workloads on Heterogeneous NPUAritra Dhar, Clément Thorens, Lara Magdalena Lazier, Lukas CavigelliS&P 2025
- LÆGIS: Pinpointing and Addressing Performance Overheads of GPU-Based Confidential ComputingYang Yang, Adwait JogISCA 2026
- Enabling Execution Assurance of Federated Learning at Untrusted ParticipantsXiaoli Zhang, Fengting Li, Zeyu Zhang, Qi Li et al.INFOCOM 2020 · 87 citations
