Vision-Language Models Can't See the Obvious
Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac, Wamiq Reyaz Para, Ankit Singh, Sanath Narayan
Abstract
We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting visually salient features that are readily apparent to humans, such as a large circle amidst a grid of smaller ones. This benchmark focuses on low-level features including color, intensity, and orientation, which are fundamental to human visual processing. Our SalBench consists of images that highlight rare, unusual, or unexpected elements within scenes, and naturally draw human attention. It comprises three novel tasks for evaluating the perceptual capabilities of LVLM: Odd-One-Out Detection, Referring Odd-One-Out, and Visual Referring Odd-One-Out. We perform a comprehensive evaluation of state-of-the-art LVLM using SalBench and our findings reveal a surprising limitation: LVLM struggle to identify seemingly obvious visual anomalies, with even the advanced GPT-4o achieving only 47.6% accuracy on such a simple task. SalBench will be an important step in measuring the capabilities of LVLM that align with the subtle definition of human attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75b59ae4-412f-4434-831d-a849e80febaeCited by top-tier papers4
- See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and ReflectionZhiheng Wu, Tong Wang, Shuning Wang, Naiming Liu et al.CVPR 2026 · 3 citations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi et al.CVPR 2026 · 2 citations
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment RewardShizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma et al.CVPR 2026 · 1 citation
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Modelstengjin Weng, Wenhao Jiang, Jingyi Wang, Ming Li et al.CVPR 2026 · 4 citations
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 1 citation
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
- SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global ThinkingSifan Li, Yujun Cai, Yiwei WangEMNLP 2025
- Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian et al.CVPR 2024
