CocktailSGD: Fine-tuning Foundation Models over 500Mbps Networks
Jue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen, Percy Liang, Christopher De Sa, Christopher Ré, Ce Zhang
Abstract
Distributed training of foundation models, especially large language models (LLMs), is communication-intensive and so has heavily relied on centralized data centers with fast interconnects. Can we train on slow networks and unlock the potential of decentralized infrastructure for foundation models? In this paper, we propose COCKTAILSGD, a novel communication-efficient training framework that combines three distinct compression techniques-random sparsification, top-K sparsification, and quantization-to achieve much greater compression than each individual technique alone. We justify the benefit of such a hybrid approach through a theoretical analysis of convergence. Empirically, we show that COCKTAILSGD achieves up to 117× compression in fine-tuning LLMs up to 20 billion parameters without hurting convergence. On a 500Mbps network, COCKTAILSGD only incurs ∼ 1.2× slowdown compared with data center networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 739e9edf-edd1-4a84-a46a-c183957c7803Cited by top-tier papers20
- Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowYixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang et al.ASPLOS 2025 · 33 citations
- Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCoZachary Charles, Gabriel Teston, Lucio M. Dery, John Keith Rush et al.NeurIPS 2025 · 29 citations
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTiancheng Chen, Ales Kubicek, Langwen Huang, Torsten HoeflerUSENIX ATC 2025 · 20 citations
- MuLoCo: Muon is a Practical Inner Optimizer for DiLoCoBenjamin Thérien, Xiaolong Huang, Aaron Defazio, Irina Rish et al.ICML 2026 · 15 citations
- Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model ParallelismSameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo et al.NeurIPS 2025 · 12 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
Related papers
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
- Fine-tuning Language Models over Slow Networks using Activation Quantization with GuaranteesJue Wang, Binhang Yuan, Luka Rimanic, Yongjun He et al.NeurIPS 2022 · 37 citations
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
- Quantized Distributed Training of Large Models with Convergence GuaranteesIlia Markov, Adrian Vladu, Qi Guo, Dan AlistarhICML 2023 · 18 citations
- Indirect Stochastic Gradient Quantization and Its Application in Distributed Deep LearningAfshin Abdi, Faramarz FekriAAAI 2020 · 5 citations
