Demystifying Supervision Data Generalization in Multimodal LMs
Xuan Qi, Luxi He, Dan Roth, Xingyu Fu
摘要
Conventional wisdom in selecting supervision data for multimodal large language models (MLLMs) is to prioritize datasets that are intuitively similar to the target task (e.g. text-rich v.s. vision-centric). However, it remains unclear how reliably such similarity translates into improved performance on the test benchmarks. In this paper, we take the first step to study the problem in MLLMs: can we predict a training data's influence on a target benchmark even before any training takes place? To answer this question, we first conduct an in-depth analysis using 14 vision-language datasets covering 7 diverse tasks. Our analysis shows that intuitive task similarity is unreliable in predicting task generalizability, and that transfer depends on the specific dataset rather than the broader task category. We propose DATAPROPHET, a training-free, simple yet effective metric based on multimodal perplexity, similarity, and data diversity. Our experiments demonstrate that the influence rankings for different supervision datasets derived from DATAPROPHET is strongly-correlated with rankings based on the actual performance increase after training, with a Kendall’s correlation coefficient of 86.0%. Moreover, we show that DATAPROPHET can help select better supervision data, achieving up to 6.9% improvement in average over uniform selection, 1.4% over SoTA training-based baseline, and 0.2% higher than oracle experiment performance-based selection. Our code and data will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- How to train data-efficient LLMsNoveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni 等ICLR 2026 · 被引用 106 次
相关 Paper
- Train-before-Test Harmonizes Language Model RankingsGuanhua Zhang, Ricardo Dominguez-Olmedo, Moritz HardtICLR 2026 · 被引用 13 次
- Ranked from Within: Ranking Large Multimodal Models Without LabelsWeijie Tu, Weijian Deng, Dylan Campbell, Yu Yao 等ICML 2025
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment QualityYuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki 等EMNLP 2025
- Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration RateQidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang 等ICCV 2025 · 被引用 8 次
- VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang 等ACL 2026
