Concept-skill Transferability-based Data Selection for Large Vision-Language Models
Jaewoo Lee, Boyang Li, Sung Ju Hwang
摘要
Instruction tuning, or supervised finetuning on extensive task-specific data, is necessary for Large Vision-Language Models (LVLMs) to generalize well across a broad range of visionlanguage (VL) tasks. However, training on large VL datasets can become prohibitively expensive. In this work, we introduce COIN-CIDE, an effective and scalable data selection technique that uses a small model as a reference model to select visual instruction tuning data for efficient finetuning of a target LVLM, focusing on diversity and transferability. Specifically, we cluster the training data using internal activations from a small model, which identifies VL concept-skill compositions needed by a target LVLM. We then sample data from these diverse clusters by considering their density and transferability, or the ability to transfer well to other concept-skill compositions. This approach ensures the diversity of these compositions, which is vital for LVLM generalization. Extensive experiments demonstrate that COINCIDE achieves superior performance and data selection efficiency against 8 strong baselines on two distinct datasets: LLaVA-1.5 and Vision-Flan. Using only 20% of the LLaVA-1.5 dataset, COINCIDE achieves performance comparable to the LVLM finetuned on the whole dataset, with 70% reduction of the wall-clock running time. On the Vision-Flan dataset, our method achieves superior results with only 16.7% of the training data. Our code is available at https://github.com/G-JWLee/COINCIDE_code .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Vision Function Layer in Multimodal LLMsCheng Shi, Yizhou Yu, Sibei YangNeurIPS 2025 · 被引用 20 次
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity OptimizationYichen Yan, Ming Zhong, Qi Zhu, Xiaoling Gu 等NeurIPS 2025 · 被引用 8 次
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment TrajectoriesNilay Naharas, Dang Nguyen, Neslihan Bulut, MohammadHossein Bateni 等ICML 2026 · 被引用 7 次
- Visual Compositional TuningXindi Wu, Hee Seung Hwang, Polina Kirichenko, Esin Tureci 等ICLR 2026 · 被引用 3 次
- Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample SelectionQian Yang, Shivam Chandhok, Oscar Mañas, Kanishk Jain 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Less is More: High-value Data Selection for Visual Instruction TuningZikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao 等ACM MM 2025 · 被引用 3 次
- Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction TuningBardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel 等CVPR 2025
- VCM: Vision Concept Modeling with Adaptive Vision Token Compression via Instruction Fine-TuningRun Luo, Renke Shan, Longze Chen, Ziqiang Liu 等NeurIPS 2025 · 被引用 1 次
- Towards Self-Refinement of Vision-Language Models with Triangular ConsistencyYunlong Deng, Guangyi Chen, Tianpei Gu, Lingjing Kong 等NeurIPS 2025 · 被引用 3 次
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
