NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model Training
Jinghan Huang, Hyungyo Kim, Nachuan Wang, Jaeyoung Kang, Hrishi Shah, Eun Kyung Lee, Minjia Zhang, Fan Lai, Nam Sung Kim
Abstract
In distributed large model training, the long communication time required to exchange large volumes of gradients and activations among GPUs dominates the training time.To reduce the communication times, lossy or lossless compression of gradients and/or activations can be employed.However, lossy compression of gradients and activations may demand more training iterations to achieve the same model accuracy and cause convergence failure, respectively.Lossless compression, on the other hand, may not reduce the volumes of gradients and activations enough to offset the significant latency associated with compression and decompression on current platforms.To address these challenges, we propose NetZIP, an algorithm/hardware co-design for in-network lossless compression of both gradients and activations.NetZIP consists of two components.(1) NetZIP-algorithm transforms gradients and activations at the bit and value levels to help lightweight standard lossless compression achieve more compression of the gradients and activations.(2) NetZIP-accelerator integrates Net-ZIP-algorithm with a lightweight lossless compression accelerator within a NIC in a bump-in-the-wire fashion to reduce the compression/decompression latency under the resource constraints.NetZIP-algorithm compresses gradients and activations 40-63 and 43-75 percentage points more, respectively, than heavy standard lossless compression for Llama-3 70B, GPT-3 175B, and Llama-3 405B.NetZIP-accelerator, implemented within FPGA-NICs and connected to commodity servers, provides orders of magnitude lower
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 583fcc3c-5d99-4d4e-9656-4dca5cdae072Cited by top-tier papers1
Ask how each one uses itRelated papers
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi et al.SIGCOMM 2026 · 1 citation
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li et al.ASPLOS 2026
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
- Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model ParallelismSameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo et al.NeurIPS 2025 · 12 citations
- Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionJaeyong Song, Jinkyu Yim, Jaewon Jung, Hongsun Jang et al.ASPLOS 2023 · 37 citations
