Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?
Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan, Joyce Chai
Abstract
Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human’s perception of reality isn’t always faithful to the physical world. This raises a key question: do VLMs have the similar kind of illusions as humans do, or do they faithfully learn to represent reality? To investigate this question, we build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusions in state-of-the-art VLMs. Our findings have shown that although the overall alignment is low, larger models are closer to human perception and more susceptible to visual illusions. Our dataset and initial findings will promote a better understanding of visual illusions in humans and machines and provide a stepping stone for future computational models that can better align humans and machines in perceiving and communicating about the shared visual world. The code and data are available at github.com/vl-illusion/dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75fa04ce-ebb3-4c72-97d0-cd9f77acce2cCited by top-tier papers9
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun et al.CVPR 2026 · 8 citations
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful IllusionsYiting Qu, Ziqing Yang, Yihan Ma, Michael Backes et al.ICCV 2025 · 6 citations
- Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsLuca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge et al.ICML 2025
- SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global ThinkingSifan Li, Yujun Cai, Yiwei WangEMNLP 2025
- The Art of Deception: Color Visual Illusions and Diffusion ModelsAlexandra Gomez-Villa, Kai Wang, C. Alejandro Párraga, Bartlomiej Twardowski et al.CVPR 2025
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
Related papers
- Evaluating Model Perception of Color Illusions in Photorealistic ScenesLingjun Mao, Zineng Tang, Alane SuhrCVPR 2025
- 3D Visual Illusion Depth EstimationChengtang Yao, Zhidan Liu, Jiaxi Zeng, Lidong Yu et al.NeurIPS 2025 · 3 citations
- VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models AlignmentLei Li, Zhihui Xie, Mukai Li, Shunian Chen et al.EMNLP 2024 · 9 citations
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsEunkyu Park, Minyeong Kim, Gunhee KimCVPR 2025
- Same or Not? Enhancing Visual Perception in Vision-Language ModelsDamiano Marsili, Aditya Mehta, Ryan Y. Lin, Georgia GkioxariCVPR 2026 · 5 citations
