From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Abstract
Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task diversity and complexity increase, EVIT faces significant challenges in resolving data conflicts. To address this limitation, we propose the Dual Low-Rank Adaptation (Dual-LoRA), a holistic-to-local framework that enhances the adapter's capacity to address data conflict through dual structural optimization. Specifically, we utilize two subspaces: a skill space for stable, holistic knowledge retention, and a rank-rectified task space that locally activates the holistic knowledge. Additionally, we introduce Visual Cue Enhancement (VCE), a multi-level local feature aggregation module designed to enrich the visionlanguage projection with local details. Our approach is both memory-and time-efficient, requiring only 1.16× the inference time of the standard LoRA method (with injection into the query and value projection layers), and just 73% of the inference time of a 4-expert LoRA-MoE. Extensive experiments on various downstream tasks and general MLLM benchmarks validate the effectiveness of our proposed methods. Our project page are publicly available at https://github.com/pengkun-jiao/Dual- LoRA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa82552b-0ae3-4508-b524-c2cf9107cbc7Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningZiqi Wang, Chang Che, Qi Wang, Yangyang Li et al.ICCV 2025 · 4 citations
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction TuningChang Che, Ziqi Wang, Pengwan Yang, Cheems Wang et al.AAAI 2026
- CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream TasksWish Suharitdamrong, Tony Alex, Muhammad Awais, Sara AtitoICML 2026
- MokA: Multimodal Low-Rank Adaptation for MLLMsYake Wei, Yu Miao, Dongzhan Zhou, Di HuNeurIPS 2025 · 8 citations
- Ensembles of Low-Rank Expert AdaptersYinghao Li, Vianne R. Gao, Chao Zhang, MohamadAli TorkamaniICLR 2025
