VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
Minkyu Kim, Sangheon Lee, Dongmin Park
摘要
The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for visionlanguage models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced reasoning required for real-world applications. In this work, we introduce VLM-SubtleBench 1 , a benchmark designed to evaluate VLMs on subtle comparative reasoning. Our benchmark covers ten difference types-Attribute, State, Emotion, Temporal, Spatial, Existence, Quantity, Quality, Viewpoint, and Action-and curate paired question-image sets reflecting these fine-grained variations. Unlike prior benchmarks restricted to natural image datasets, our benchmark spans diverse domains, including industrial, aerial, and medical imagery. Through extensive evaluation of both proprietary and open-source VLMs, we reveal systematic gaps between model and human performance across difference types and domains, and provide controlled analyses highlighting where VLMs' reasoning sharply deteriorates. Together, our benchmark and findings establish a foundation for advancing VLMs toward human-level comparative reasoning. Recently, vision-language models (VLMs) have shown remarkable progress toward artificial general intelligence (AGI), demonstrating promising results in various tasks, such as visual question answering (VQA) and scene description (Zhang et al., 2024 ). Yet, most progress has primarily centered on single visual inputs, e.g., an image or a video, while comparative tasks that require comparison over * Equal contribution. † Work done during an internship at KRAFTON.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen 等ICLR 2024 · 被引用 258 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
相关 Paper
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary DomainSong Jin, Juntian Zhang, Xun Zhang, Zeying Tian 等ACL 2026 · 被引用 1 次
- Vision-Language Models Do Not Understand NegationKumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li 等CVPR 2025
- Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning BenchmarksMiao Jing, Mengting Jia, Junling Lin, Zhongxia Shen 等ICLR 2026 · 被引用 4 次
- SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik 等ICLR 2026 · 被引用 13 次
