LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma
摘要
Multimodal Large Language Models (MLLMs) enhance visual tasks by integrating visual representations into large language models (LLMs). The textual modality, inherited from LLMs, enables instruction following and in-context learning, while the visual modality boosts downstream task performance through rich semantic content, spatial information, and grounding capabilities. These modalities work synergistically across various visual tasks. Our research reveals a persistent imbalance between these modalities, with text often dominating output generation during visual instruction tuning, regardless of using full or parameter-efficient fine-tuning (PEFT). We found that re-balancing these modalities can significantly reduce trainable parameters, inspiring further optimization of visual instruction tuning. To this end, we introduce Modality Linear Representation-Steering (MoReS), which re-balances intrinsic modalities by steering visual representations through linear transformations in the visual subspace across each model layer. We validated our approach by developing LLaVA Steering, a suite of models using MoReS. Results show that LLaVA Steering requires, on average, 500 times fewer trainable parameters than LoRA while maintaining comparable performance across three visual benchmarks and eight visual question-answering tasks. Finally, we introduce the LLaVA Steering Factory, a platform that enables rapid customization of MLLMs with a component-based architecture, seamlessly integrating state-of-theart models and evaluating intrinsic modality imbalance. This open-source project facilitates a deeper understanding of MLLMs within the research community. Code is available at https://github.com/bibisbar/LLaVA-Steering .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language ModelsXinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen 等CVPR 2026 · 被引用 30 次
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang 等CVPR 2026 · 被引用 16 次
- Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image RetrievalZhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang 等CVPR 2026 · 被引用 16 次
- INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li 等AAAI 2026 · 被引用 12 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang 等NeurIPS 2024 · 被引用 24 次
- Explore How to Inject Beneficial Noise in MLLMsRuishu Zhu, Sida Huang, Ziheng Jiao, Hongyuan ZhangAAAI 2026 · 被引用 7 次
- MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMsWayner Barrios, Andrés Villa, Juan Leon Alcazar, SouYoung Jin 等ICML 2026
- Parameter-Efficient Adaptation for MLLMs via Implicit Modality DecompositionMingfang Zhang, Yunhong Wang, Lu Wang, Jiaxin ChenCVPR 2026
- LRM-LLaVA: Overcoming the Modality Gap of Multilingual Large Language-Vision Model for Low-Resource LanguagesJunchen Li, Qing Yang, Bojian Jiang, Shaolin Zhu 等AAAI 2025 · 被引用 3 次
