Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices
Yuhao Chen, Yuxuan Yan, Shuowei Ge, Yuyang Qin, Yue Zheng, Qianqian Yang, Shibo He, Zhiguo Shi, Jiming Chen, Yuanchao Shu
Abstract
Large language models (LLMs) have emerged as a cornerstone for advancing AI technologies. It revolutionizes the way we interact with devices, websites, and information, and paves the way for the development of highly intuitive and capable virtual assistants. Training of today's LLMs happens in cloud data centers due to the requirement of enormous data and a significant amount of computing power. Despite extensive research in mobile edge computing, fine-tuning pre-trained LLMs using resource-constrained devices like commodity smartphones remains highly under-explored. In this paper, we propose Confidant, a practical collaborative training framework that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. To this end, Confidant partitions an LLM into several sub-models, allowing each of them to fit in the memory of a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. In specific, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. To ensure resilient distributed training, a hybrid fault tolerance mechanism is devised to proactively manage potential device and network failures. We fully implemented Confidant in C++/Python, and built a cross-framework adapter, enabling collaborative training on a variety of mobile platforms. Experimental results show that Confidant excels in achieving computation-, memory-efficient, and robust customization of LLMs - it manages to train state-of-the-art billion-sized LLMs including BERT, GPT-2, Phi2, and LLaMA3, and fine-tunes Phi2-2.7B on Alpaca in just 40.1 hours using three consumer-grade mobile devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45c60426-672a-422a-9391-ac3a5fc8ac77Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- Distributed Inference and Fine-tuning of Large Language Models Over The InternetAlexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk et al.NeurIPS 2023 · 108 citations
- Titanic: Towards Production Federated Learning with Large Language ModelsNingxin Su, Chenghao Hu, Baochun Li, Bo LiINFOCOM 2024 · 30 citations
- Federated Adaptive Fine-Tuning of Large Language Models with Heterogeneous Quantization and LoRAZhidong Gao, Zhenxiao Zhang, Yuanxiong Guo, Yanmin GongINFOCOM 2025 · 12 citations
- A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge SystemsYuze Liu, Yunhan Wang, Tiehua Zhang, Zhishu Shen et al.WWW 2026 · 1 citation
- FedDynMask: Efficient Federated Fine-Tuning for Edge LLMs via Dynamic Sparse MaskingYan Wang, Ziyi Gao, Yida Zhang, Rui WangINFOCOM 2026
