Vision Language Models are Biased
An Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang, Anh Totti Nguyen, Daeyoung Kim
摘要
Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may notoriously sway their outputs toward wrong or biased answers. In this work, we test how the knowledge of popular subjects hurts the accuracy of vision language models (VLMs) on standard, objective visual tasks of counting and identification. We find that stateof-the-art VLMs are strongly biased (e.g., unable to recognize that a 4th stripe has been added to a 3-stripe Adidas logo), scoring an average of 17.05% accuracy in counting (e.g., counting stripes in an Adidas-like logo) across 7 diverse domains spanning animals, logos, chess, game boards, optical illusions, and patterned grids. Removing image backgrounds nearly doubles accuracy (by 21.09 points), revealing that background visual cues trigger these biased responses. Further analysis of VLMs' reasoning patterns shows that counting accuracy initially rises with thinking tokens, reaching ∼40%, before declining with model overthinking. Our work presents an interesting failure mode in VLMs and a human-supervised automated framework for testing VLM biases. Code and data are available at: vlmsarebiased.github.io. * Equal contribution. † Equal advising.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Understanding Language Prior of LVLMs by Contrasting Chain-of-EmbeddingLin Long, Changdae Oh, Seongheon Park, Sharon LiICLR 2026 · 被引用 14 次
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn EditingTianyu Chen, Yasi Zhang, Zhi Zhang, Peiyu Yu 等ICLR 2026 · 被引用 11 次
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun 等CVPR 2026 · 被引用 8 次
- Do 3D Large Language Models Really Understand 3D Spatial Relationships?Xianzheng Ma, Tao Sun, Shuai Chen, Yash Bhalgat 等ICLR 2026 · 被引用 7 次
- FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMsZhihan Yin, Jianxin Liang, Yueqian Wang, Yifeng Yao 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper21
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank 等ACL 2022 · 被引用 147 次
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
相关 Paper
- A Peek into Token Bias: Large Language Models Are Not Yet Genuine ReasonersBowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang 等EMNLP 2024 · 被引用 27 次
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 被引用 6 次
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language ModelsChahat Raj, Bowen Wei, Aylin Caliskan, Antonios Anastasopoulos 等ACL 2026 · 被引用 3 次
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language ModelsMd. Atabuzzaman, Ali Asgarov, Christopher ThomasEMNLP 2025
