MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering
Jingjing Jiang, Nanning Zheng
Abstract
Recently, finetuning pretrained vision-language models (VLMs) has been a prevailing paradigm for achieving stateof-the-art performance in VQA. However, as VLMs scale, it becomes computationally expensive, storage inefficient, and prone to overfitting when tuning full model parameters for a specific task in low-resource settings. Although current parameter-efficient tuning methods dramatically reduce the number of tunable parameters, there still exists a significant performance gap with full finetuning. In this paper, we propose MixPHM, a redundancy-aware parameterefficient tuning method that outperforms full finetuning in low-resource VQA. Specifically, MixPHM is a lightweight module implemented by multiple PHM-experts in a mixtureof-experts manner. To reduce parameter redundancy, we reparameterize expert weights in a low-rank subspace and share part of the weights inside and across MixPHM. Moreover, based on our quantitative analysis of representation redundancy, we propose Redundancy Regularization, which facilitates MixPHM to reduce task-irrelevant redundancy while promoting task-relevant correlation. Experiments conducted on VQA v2, GQA, and OK-VQA with different low-resource settings show that our MixPHM outperforms state-of-the-art parameter-efficient methods and is the only one consistently surpassing full finetuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9084918-ade7-4bb0-badb-8686ad54a226Cited by top-tier papers3
- Improving Zero-shot Visual Question Answering via Large Language Models with Reasoning Question PromptsYunshi Lan, Xiang Li, Xin Liu, Yang Li et al.ACM MM 2023 · 29 citations
- Medusa: A Framework for Collaborative Development of Foundation Models with Automated Parameter Ownership AssignmentDezhi Ran, Yuan Cao, Yuzhe Guo, Yuetong Li et al.FSE 2025 · 1 citation
- Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model VariantsChong Yu, Tao Chen, Zhongxue GanCVPR 2025
Builds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Self-PT: Adaptive Self-Prompt Tuning for Low-Resource Visual Question AnsweringBowen Yuan, Sisi You, Bing-Kun BaoACM MM 2023 · 5 citations
- Revisit Visual Prompt Tuning: The Expressiveness of Prompt ExpertsMinh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen et al.ICLR 2026 · 6 citations
- Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsZihan Wang, Deli Chen, Damai Dai, Runxin Xu et al.EMNLP 2024 · 2 citations
- Parameters as Experts: Adapting Vision Models with Dynamic Parameter RoutingMeng Lou, Stanley Yu, Yizhou YuICML 2026 · 1 citation
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
