Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
Rohit Sinha, Aditya Sanjiv Kanade, Sai Srinivas Kancheti, Vineeth N. Balasubramanian, Tanuja Ganu
Abstract
Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce Mind's Eye, a multiple-choice benchmark of eight visuo-cognitive tasks inspired by classic human intelligence tests and organized under a novel A-R-T taxonomy: Abstraction, Relation, and Transformation. The tasks probe core processes of fluid intelligence such as pattern induction, analogical relation mapping, and mental transformation. We evaluate a diverse suite of closed-source and open-source MLLMs and compare their performance with human participants. Humans achieve 80% accuracy, while top performing MLLMs remain below 50%. Error analysis reveals failures in: (i) visual attention allocation, (ii) internal perceptual manipulation, and (iii) weak abstraction of underlying visual concepts. Our findings suggest that current MLLMs exhibit limited visuospatial reasoning capabilities, when compared with human participants, highlighting the need for more cognitively grounded evaluation frameworks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6cd88b5-40ec-4554-817f-4f8eaa97a7f7Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
Related papers
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMsSiting Wang, Minnan Pei, Luoyang Sun, Cheng Deng et al.ICLR 2026 · 8 citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual SimulationsLinjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma et al.ICLR 2026 · 22 citations
- KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal ModelsEunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong et al.ICLR 2025
