Explore How to Inject Beneficial Noise in MLLMs
Ruishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan Zhang
Abstract
Multimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore crossmodal heterogeneity, limiting their full potential. In this work, we propose a novel fine-tuning strategy by injecting beneficial random noise, which outperforms previous methods and even surpasses full fine-tuning, with minimal additional parameters. The proposed Multimodal Noise Generator (MuNG) enables efficient modality fine-tuning by injecting customized noise into the frozen MLLMs. Specifically, we reformulate the reasoning process of MLLMs from a variational inference perspective, upon which we design a multimodal noise generator that dynamically analyzes crossmodal relationships in image-text pairs to generate taskadaptive beneficial noise. Injecting this type of noise into the MLLMs effectively suppresses irrelevant semantic components, leading to significantly improved cross-modal representation alignment and enhanced performance on downstream tasks. Experiments on two mainstream MLLMs, QwenVL and LLaVA, demonstrate that our method surpasses full-parameter fine-tuning and other existing fine-tuning approaches, while requiring adjustments to only about 1 ∼ 2% additional parameters. The relevant code is uploaded in the supplementary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c09b3e04-4d87-44c6-813a-2b66862a4793Cited by top-tier papers1
Ask how each one uses itBuilds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
Related papers
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient AdaptationJingyi Xie, Jintao Yang, Zhunchen Luo, Yunbo Cao et al.CVPR 2025
- MokA: Multimodal Low-Rank Adaptation for MLLMsYake Wei, Yu Miao, Dongzhan Zhou, Di HuNeurIPS 2025 · 8 citations
- Differential Fine-Tuning Large Language Models Towards Better Diverse Reasoning AbilitiesXiaosong Yuan, Chen Shen, Shaotian Yan, kaiyuan liu et al.ICLR 2026 · 6 citations
- CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language ModelsNan Zhou, Huiqun Wang, Yaoyan Zheng, Di HuangCVPR 2026
