Cloud-Device Collaborative Learning for Multimodal Large Language Models
Guanqun Wang, Jiaming Liu, Chenxuan Li, Yuan Zhang, Junpeng Ma, Xinyu Wei, Kevin Zhang, Maurice Chong, Renrui Zhang, Yijiang Liu, Shanghang Zhang
摘要
The burgeoning field of Multimodal Large Language Models (MLLMs) has exhibited remarkable performance in diverse tasks such as captioning, commonsense reasoning, and visual scene understanding. However, the deployment of these large-scale MLLMs on client devices is hindered by their extensive model parameters, leading to a notable de-cline in generalization capabilities when these models are compressed for device deployment. Addressing this chal-lenge, we introduce a Cloud-Device Collaborative Contin-ual Adaptation framework, designed to enhance the performance of compressed, device-deployed MLLMs by lever-aging the robust capabilities of cloud-based, larger-scale MLLMs. Our framework is structured into three key components: a device-to-cloud uplink for efficient data transmission, cloud-based knowledge adaptation, and an optimized cloud-to-device downlink for model deployment. In the up-link phase, we employ an Uncertainty-guided Token Sam-pling (UTS) strategy to effectively filter out-of-distribution tokens, thereby reducing transmission costs and improving training efficiency. On the cloud side, we propose Adapter-based Knowledge Distillation (AKD) method to transfer refined knowledge from large-scale to compressed, pocket-size MLLMs. Furthermore, we propose a Dynamic Weight update Compression (DWC) strategy for the down-link, which adaptively selects and quantizes updated weight parameters, enhancing transmission efficiency and reducing the representational disparity between cloud and de-vice models. Extensive experiments on several multimodal benchmarks demonstrate the superiority of our proposed framework over prior Knowledge Distillation and device-cloud collaboration methods. Notably, we also validate the feasibility of our approach to real-world experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image CompositionXinyu Wei, Kangrui Cen, Hongyang Wei, Zhen Guo 等CVPR 2026 · 被引用 10 次
- Enabling Real-Time Inference in Online Continual Learning via Device-Cloud CollaborationHaibo Liu, Chen Gong, Zhenzhe Zheng, Shengzhong Liu 等WWW 2025 · 被引用 10 次
- A Population-to-individual Tuning Framework for Adapting Pretrained LM to On-device User Intent PredictionJiahui Gong, Jingtao Ding, Fanjin Meng, Guilong Chen 等KDD 2024 · 被引用 7 次
- LSRP: A Leader-Subordinate Retrieval Framework for Privacy-Preserving Cloud-Device CollaborationYingyi Zhang, Pengyue Jia, Xianneng Li, Derong Xu 等KDD 2025 · 被引用 2 次
- CIAR: Interval-based Collaborative Decoding for Image Generation AccelerationKeming Ye, Zhou Zhao, Fan Wu, Shengyu ZhangICLR 2026
它引用的顶会 Paper22
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
相关 Paper
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He 等ICCV 2025 · 被引用 9 次
- Cloud-Device Collaborative Adaptation to Continual Changing Environments in the Real-WorldYulu Gan, Mingjie Pan, Rongyu Zhang, Zijian Ling 等CVPR 2023
- AVAM: A Universal Training-Free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-Image Question AnsweringKang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan 等ICCV 2025 · 被引用 1 次
- Bridging Compressed Image Latents and Multimodal Large Language ModelsChia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen 等ICLR 2025
- Efficient Multimodal Large Language Model via Dynamic KV Cache QuantizationJiahao Fan, Chien-Ming ChenAAAI 2026
