Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?
Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan, Joyce Chai
摘要
Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human’s perception of reality isn’t always faithful to the physical world. This raises a key question: do VLMs have the similar kind of illusions as humans do, or do they faithfully learn to represent reality? To investigate this question, we build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusions in state-of-the-art VLMs. Our findings have shown that although the overall alignment is low, larger models are closer to human perception and more susceptible to visual illusions. Our dataset and initial findings will promote a better understanding of visual illusions in humans and machines and provide a stepping stone for future computational models that can better align humans and machines in perceiving and communicating about the shared visual world. The code and data are available at github.com/vl-illusion/dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun 等CVPR 2026 · 被引用 8 次
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful IllusionsYiting Qu, Ziqing Yang, Yihan Ma, Michael Backes 等ICCV 2025 · 被引用 6 次
- Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsLuca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge 等ICML 2025
- SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global ThinkingSifan Li, Yujun Cai, Yiwei WangEMNLP 2025
- The Art of Deception: Color Visual Illusions and Diffusion ModelsAlexandra Gomez-Villa, Kai Wang, C. Alejandro Párraga, Bartlomiej Twardowski 等CVPR 2025
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
相关 Paper
- Evaluating Model Perception of Color Illusions in Photorealistic ScenesLingjun Mao, Zineng Tang, Alane SuhrCVPR 2025
- 3D Visual Illusion Depth EstimationChengtang Yao, Zhidan Liu, Jiaxi Zeng, Lidong Yu 等NeurIPS 2025 · 被引用 3 次
- VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models AlignmentLei Li, Zhihui Xie, Mukai Li, Shunian Chen 等EMNLP 2024 · 被引用 9 次
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsEunkyu Park, Minyeong Kim, Gunhee KimCVPR 2025
- Same or Not? Enhancing Visual Perception in Vision-Language ModelsDamiano Marsili, Aditya Mehta, Ryan Y. Lin, Georgia GkioxariCVPR 2026 · 被引用 5 次
