Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems
Kai Huang, Xiangyu Yin, Heng Huang, Wei Gao
摘要
Multimodal reasoning by LLMs is critical to autonomous mobile systems, but the growing diversity of input data modalities prevents incorporating all modalities into LLMs. Instead, only the useful modalities should be adaptively involved at runtime, based on the current environmental contexts and task requirements. Existing work on runtime modality adaptation uses fixed connections between data encoders and LLM's input layer, but results in high training costs and ineffective cross-modal interaction. In this paper, we present MPnP, a new modality adaptation technique that connects data encoders to a flexible set of last LLM blocks and makes such latent connections fully trainable at runtime. Evaluation results show that MPnP has high compute and data efficiency, with 3.7× FLOPs reduction and 30% memory usage reduction compared to best baselines. It requires only few hundreds of training samples at runtime, and completes modality adaptation within few minutes on weak devices.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- When Device Delays Meet Data Heterogeneity in Federated AIoT ApplicationsHaoming Wang, Wei GaoMobiCom 2025 · 被引用 2 次
- PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video GenerationQiyao Xue, Xiangyu Yin, Boyuan Yang, Wei GaoCVPR 2025
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui 等AAAI 2026
相关 Paper
- MokA: Multimodal Low-Rank Adaptation for MLLMsYake Wei, Yu Miao, Dongzhan Zhou, Di HuNeurIPS 2025 · 被引用 8 次
- TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance TransferCheng Fang, Zimu Zhou, Ke Ma, Bin GuoCVPR 2026
- Scaling Laws for Native Multimodal ModelsMustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord 等ICCV 2025 · 被引用 4 次
- Learning to Inference Adaptively for Multimodal Large Language ModelsZhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi 等ICCV 2025 · 被引用 4 次
- AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionRameswar Panda, Chun-Fu (Richard) Chen, Quanfu Fan, Ximeng Sun 等ICCV 2021 · 被引用 65 次
