Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems
Kai Huang, Xiangyu Yin, Heng Huang, Wei Gao
Abstract
Multimodal reasoning by LLMs is critical to autonomous mobile systems, but the growing diversity of input data modalities prevents incorporating all modalities into LLMs. Instead, only the useful modalities should be adaptively involved at runtime, based on the current environmental contexts and task requirements. Existing work on runtime modality adaptation uses fixed connections between data encoders and LLM's input layer, but results in high training costs and ineffective cross-modal interaction. In this paper, we present MPnP, a new modality adaptation technique that connects data encoders to a flexible set of last LLM blocks and makes such latent connections fully trainable at runtime. Evaluation results show that MPnP has high compute and data efficiency, with 3.7× FLOPs reduction and 30% memory usage reduction compared to best baselines. It requires only few hundreds of training samples at runtime, and completes modality adaptation within few minutes on weak devices.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7b181586-fefa-46bf-808c-23e7c71efda1Cited by top-tier papers3
- When Device Delays Meet Data Heterogeneity in Federated AIoT ApplicationsHaoming Wang, Wei GaoMobiCom 2025 · 2 citations
- PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video GenerationQiyao Xue, Xiangyu Yin, Boyuan Yang, Wei GaoCVPR 2025
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui et al.AAAI 2026
Related papers
- MokA: Multimodal Low-Rank Adaptation for MLLMsYake Wei, Yu Miao, Dongzhan Zhou, Di HuNeurIPS 2025 · 8 citations
- TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance TransferCheng Fang, Zimu Zhou, Ke Ma, Bin GuoCVPR 2026
- Scaling Laws for Native Multimodal ModelsMustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord et al.ICCV 2025 · 4 citations
- Learning to Inference Adaptively for Multimodal Large Language ModelsZhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi et al.ICCV 2025 · 4 citations
- AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionRameswar Panda, Chun-Fu (Richard) Chen, Quanfu Fan, Ximeng Sun et al.ICCV 2021 · 65 citations
