Less is More: High-value Data Selection for Visual Instruction Tuning
Zikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao, Yaliang Li, Ji-Rong Wen
Abstract
Visual instruction tuning is the key to building large vision language models (LVLMs), which can greatly improve the task generalization and solving capabilities by learning a mixture of instruction data from diverse visual tasks. Previous work mostly collects multiple existing visual instruction datasets via heuristic ways for training (even more than a million instructions), which may introduce data redundancy and enlarge the training cost. To investigate this issue, we conduct a series of empirical studies, which reveal a significant redundancy within the visual instruction datasets, and show that greatly reducing the amount of instructions from several tasks even do not affect the performance. Based on the findings, we propose a high-value data selection approach TIVE, to eliminate redundancy within the visual instruction data and reduce the training cost. In TIVE, we first estimate the instance influence score on its corresponding task, and the task difficulty score, based on the gradient-based influence functions. Then, we leverage the two kinds of scores to determine the task proportion within the selected visual instruction subset, and select high-value instances for each task, respectively. Experiments on various LVLMs show that our approach using only about 15% data can achieve comparable average performance to the full-data fine-tuned model across eight benchmarks, even surpassing it on four of the benchmarks. Our code and data will be publicly released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b94a6d69-d6a1-4ef4-a725-1b7b564dc19bCited by top-tier papers12
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity OptimizationYichen Yan, Ming Zhong, Qi Zhu, Xiaoling Gu et al.NeurIPS 2025 · 8 citations
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment TrajectoriesNilay Naharas, Dang Nguyen, Neslihan Bulut, MohammadHossein Bateni et al.ICML 2026 · 7 citations
- Visual Compositional TuningXindi Wu, Hee Seung Hwang, Polina Kirichenko, Esin Tureci et al.ICLR 2026 · 3 citations
- Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample SelectionQian Yang, Shivam Chandhok, Oscar Mañas, Kanishk Jain et al.CVPR 2026 · 3 citations
- Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal ModelsXu Yang, Chen Liu, Ying WeiNeurIPS 2025 · 2 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction TuningBardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel et al.CVPR 2025
- Concept-skill Transferability-based Data Selection for Large Vision-Language ModelsJaewoo Lee, Boyang Li, Sung Ju HwangEMNLP 2024 · 1 citation
- Mastering Collaborative Multi-Modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessQifan Yu, Zhebei Shen, Zhongqi Yue, Yang Wu et al.ICCV 2025 · 1 citation
- Importance-Aware Data Selection for Efficient LLM Instruction TuningTingyu Jiang, Shen Li, Yiyao Song, Lan Zhang et al.AAAI 2026 · 5 citations
- What Makes Good Instruction-Tuning Data? An In-Context Learning PerspectiveGuangzeng Han, Xiaolei HuangACL 2026 · 1 citation
