Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Yikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen, Qianyu Zhou, Yunhai Tong, Shunping Ji, Jiangning Zhang, Lu Qi, Xiangtai Li
Abstract
Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential in computer vision. Our research reveals that the matching capabilities in recent MLLMs still exhibit systematic shortcomings, even with current strong MLLMs models, GPT-4o. In particular, we construct a Multimodal Visual Matching (MMVM) benchmark to fairly benchmark over 30 different MLLMs. The MMVM benchmark is built from 15 open-source datasets and Internet videos with manual annotation. We categorize the data samples of MMVM benchmark into eight aspects based on the required cues and capabilities to more comprehensively evaluate and analyze current MLLMs. In addition, we have designed an automatic annotation pipeline to generate the MMVM SFT dataset, including 220K visual matching data with reasoning annotation. To our knowledge, this is the first visual corresponding dataset and benchmark for the MLLM community. Finally, we present CoLVA, a novel contrastive MLLM with two novel technical designs: fine-grained vision expert with object-level contrastive learning and instruction augmentation strategy. The former learns instance discriminative tokens, while the latter further improves instruction following ability. CoLVA-InternVL2-4B achieves an overall accuracy (OA) of 49.80% on the MMVM benchmark, surpassing GPT-4o and the best open-source MLLM, Qwen2VL-72B, by 7.15% and 11.72% OA, respectively. These results demonstrate the effectiveness of our MMVM SFT dataset and our novel technical designs. Code, benchmark, dataset, and models will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ead8aa64-d7c3-4199-ab3c-6a4ce7655af2Cited by top-tier papers2
- SAMTok: Representing Any Mask with Two WordsYikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu et al.CVPR 2026 · 10 citations
- MiCo: Multi-image Contrast for Reinforcement Visual ReasoningXi Chen, Mingkang Zhu, Shaoteng Liu, Xiaoyang Wu et al.NeurIPS 2025 · 8 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Img-Diff: Contrastive Data Synthesis for Multimodal Large Language ModelsQirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding et al.CVPR 2025
- Hallucination Augmented Contrastive Learning for Multimodal Large Language ModelChaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen et al.CVPR 2024
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 1 citation
- Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMsWenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang et al.ACL 2025
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma et al.CVPR 2024 · 111 citations
