Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language Models
Yapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang, Zicong Hong, Song Guo, Yue Yu
2026Year
Abstract
Large language models (LLMs) have driven significant progress in natural language processing, yet their training and fine-tuning remain limited by memory constraints, particularly the substantial memory footprints of optimizer states. Existing solutions address this challenge by offloading optimizer states and update tasks to the CPU, but this often leads to increased GPU idleness due to the CPU's limited computational capabilities, especially in pipeline parallelism.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5b427ea1-6d2f-4245-a6a5-c6e6a5b0f123Related papers
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo et al.ACL 2024 · 61 citations
- Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse ProjectorsSiyuan Chen, Zhuofeng Wang, Zelong Guan, Yudong Liu et al.AAAI 2025 · 3 citations
- ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM TrainingAdel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano et al.NeurIPS 2025 · 5 citations
- LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM SystemsWei Liu, Yongchao He, Bohan Zhao, Hongyi Wang et al.SOSP 2026
- Efficient Long Context Fine-tuning with Chunk FlowXiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang et al.ICML 2025
