NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Training
Jinghan Huang, Hyungyo Kim, Nachuan Wang, Jaeyoung Kang, Hrishi Shah, Eun Kyung Lee, Minjia Zhang, Fan Lai, Nam Sung Kim
摘要
In distributed large model training, the long communication time required to exchange large volumes of gradients and activations among GPUs dominates the training time.To reduce the communication times, lossy or lossless compression of gradients and/or activations can be employed.However, lossy compression of gradients and activations may demand more training iterations to achieve the same model accuracy and cause convergence failure, respectively.Lossless compression, on the other hand, may not reduce the volumes of gradients and activations enough to offset the significant latency associated with compression and decompression on current platforms.To address these challenges, we propose NetZIP, an algorithm/hardware co-design for in-network lossless compression of both gradients and activations.NetZIP consists of two components.(1) NetZIP-algorithm transforms gradients and activations at the bit and value levels to help lightweight standard lossless compression achieve more compression of the gradients and activations.(2) NetZIP-accelerator integrates Net-ZIP-algorithm with a lightweight lossless compression accelerator within a NIC in a bump-in-the-wire fashion to reduce the compression/decompression latency under the resource constraints.NetZIP-algorithm compresses gradients and activations 40-63 and 43-75 percentage points more, respectively, than heavy standard lossless compression for Llama-3 70B, GPT-3 175B, and Llama-3 405B.NetZIP-accelerator, implemented within FPGA-NICs and connected to commodity servers, provides orders of magnitude lower
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi 等SIGCOMM 2026 · 被引用 1 次
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li 等ASPLOS 2026
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
- Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model ParallelismSameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo 等NeurIPS 2025 · 被引用 12 次
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang 等ASPLOS 2023 · 被引用 37 次
