Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Juncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao, Wei Ji, Wenqiao Zhang, Tat-Seng Chua, Siliang Tang, Hanwang Zhang, Yueting Zhuang
Abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have been utilizing Visual Prompt Generators (VPGs) to convert visual features into tokens that LLMs can recognize. This is achieved by training the VPGs on millions of image-caption pairs, where the VPG-generated tokens of images are fed into a frozen LLM to generate the corresponding captions. However, this image-captioning based training objective inherently biases the VPG to concentrate solely on the primary visual contents sufficient for caption generation, often neglecting other visual details. This shortcoming results in MLLMs' underperformance in comprehending demonstrative instructions consisting of multiple, interleaved, and multimodal instructions that demonstrate the required context to complete a task. To address this issue, we introduce a generic and lightweight Visual Prompt Generator Complete module (VPG-C), which can infer and complete the missing details essential for comprehending demonstrative instructions. Further, we propose a synthetic discriminative training strategy to fine-tune VPG-C, eliminating the need for supervised demonstrative instructions. As for evaluation, we build DEMON, a comprehensive benchmark for demonstrative instruction understanding. Synthetically trained with the proposed strategy, VPG-C achieves significantly stronger zero-shot performance across all tasks of DEMON. Further evaluation on the MME and OwlEval benchmarks also demonstrate the superiority of VPG-C. Our benchmark, code, and pre-trained models are available at https://github.com/DCDmllm/Cheetah.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7acaa53b-0173-45e1-9ba4-9cdf096c5921Cited by top-tier papers33
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language TransformersDachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang et al.ICML 2024 · 46 citations
- Auto-Encoding Morph-Tokens for Multimodal LLMKaihang Pan, Siliang Tang, Juncheng Li, Zhaoyu Fan et al.ICML 2024 · 36 citations
- WorldGPT: Empowering LLM as Multimodal World ModelZhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li et al.ACM MM 2024 · 35 citations
- Data Shunt: Collaboration of Small and Large Models for Lower Costs and Better PerformanceDong Chen, Yueting Zhuang, Shuo Zhang, Jinfeng Liu et al.AAAI 2024 · 32 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal ModelsKuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang et al.EMNLP 2025
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 4 citations
- PGT: Procedurally Generated Tasks for improving visual grounding in MLLMsRim Assouel, Amir Bar, Michal Drozdzal, Adriana Romero-SorianoICML 2026
- ProgressLM: Towards Progress Reasoning in Vision-Language ModelsJianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu et al.ACL 2026 · 7 citations
