Learning to Instruct for Visual Instruction Tuning
Zhihan Zhou, Feng Hong, Jiaan Luo, Yushi Ye, Jiangchao Yao, Dongsheng Li, Bo Han, Ya Zhang, Yanfeng Wang
摘要
We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning, potentially degrading performance. This gap arises from an overemphasis on instruction-following abilities, while neglecting the proactive understanding of visual information. Inspired by this, L2T adopts a simple yet effective approach by incorporating the loss function into both the instruction and response sequences. It seamlessly expands the training data, and regularizes the MLLMs from overly relying on language priors. Based on this merit, L2T achieves a significant relative improvement of up to 9% on comprehensive multimodal benchmarks, requiring no additional training data and incurring negligible computational overhead. Surprisingly, L2T attains exceptional fundamental visual capabilities, yielding up to an 18% improvement in captioning performance, while simultaneously alleviating hallucination in MLLMs. Github code: https://github.com/Feng-Hong/L2T.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Visual Instruction Bottleneck TuningChangdae Oh, Jiatong Li, Shawn Im, Sharon LiNeurIPS 2025 · 被引用 7 次
- Differential-Informed Sample Selection Accelerates Multimodal Contrastive LearningZihua Zhao, Feng Hong, Mengxi Chen, Pengyi Chen 等ICCV 2025
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- MetaMorph: Multimodal Understanding and Generation via Instruction TuningShengbang Tong, David Fan, Jiachen Zhu, Yunyang Xiong 等ICCV 2025 · 被引用 14 次
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningTaowen Wang, Yiyang Liu, James Liang, Junhan Zhao 等EMNLP 2024 · 被引用 31 次
- LLaMA-Excitor: General Instruction Tuning via Indirect Feature InteractionBo Zou, Chao Yang, Yu Qiao, Chengbin Quan 等CVPR 2024 · 被引用 5 次
