Cloud-Device Collaborative Learning for Multimodal Large Language Models
Guanqun Wang, Jiaming Liu, Chenxuan Li, Yuan Zhang, Junpeng Ma, Xinyu Wei, Kevin Zhang, Maurice Chong, Renrui Zhang, Yijiang Liu, Shanghang Zhang
Abstract
The burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene understanding. However, the deployment of these large-scale MLLMs on client devices is hindered by their extensive model parameters, leading to a notable de-cline in generalization capabilities when these models are compressed for device deployment. Addressing this chal-lenge, we introduce a Cloud-Device Collaborative Contin-ual Adaptation framework, designed to enhance the performance of compressed, device-deployed MLLMs by lever-aging the robust capabilities of cloud-based, larger-scale MLLMs. Our framework is structured into three key components: a device-to-cloud uplink for efficient data transmission, cloud-based knowledge adaptation, and an optimized cloud-to-device downlink for model deployment. In the up-link phase, we employ an Uncertainty-guided Token Sam-pling (UTS) strategy to effectively filter out-of-distribution tokens, thereby reducing transmission costs and improving training efficiency. On the cloud side, we propose Adapter-based Knowledge Distillation (AKD) method to transfer refined knowledge from large-scale to compressed, pocket-size MLLMs. Furthermore, we propose a Dynamic Weight update Compression (DWC) strategy for the down-link, which adaptively selects and quantizes updated weight parameters, enhancing transmission efficiency and reducing the representational disparity between cloud and de-vice models. Extensive experiments on several multimodal benchmarks demonstrate the superiority of our proposed framework over prior Knowledge Distillation and device-cloud collaboration methods. Notably, we also validate the feasibility of our approach to real-world experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c70850b-d05b-4ab8-beb2-e69353a4706bCited by top-tier papers7
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo et al.CVPR 2026 · 10 citations
- Enabling Real-Time Inference in Online Continual Learning via Device-Cloud CollaborationHaibo Liu, Chen Gong, Zhenzhe Zheng, Shengzhong Liu et al.WWW 2025 · 10 citations
- A Population-to-individual Tuning Framework for Adapting Pretrained LM to On-device User Intent PredictionJiahui Gong, Jingtao Ding, Fanjin Meng, Guilong Chen et al.KDD 2024 · 7 citations
- LSRP: A Leader-Subordinate Retrieval Framework for Privacy-Preserving Cloud-Device CollaborationYingyi Zhang, Pengyue Jia, Xianneng Li, Derong Xu et al.KDD 2025 · 2 citations
- CIAR: Interval-based Collaborative Decoding for Image Generation AccelerationKeming Ye, Zhou Zhao, Fan Wu, Shengyu ZhangICLR 2026
Builds on22
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
Related papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- Cloud-Device Collaborative Adaptation to Continual Changing Environments in the Real-WorldYulu Gan, Mingjie Pan, Rongyu Zhang, Zijian Ling et al.CVPR 2023
- AVAM: A Universal Training-Free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-Image Question AnsweringKang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan et al.ICCV 2025 · 1 citation
- Bridging Compressed Image Latents and Multimodal Large Language ModelsChia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen et al.ICLR 2025
- Efficient Multimodal Large Language Model via Dynamic KV Cache QuantizationJiahao Fan, Chien-Ming ChenAAAI 2026
