KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Eunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong, Yutong Bai, Shiry Ginosar, Alison Gopnik, Kate Saenko
Abstract
This paper investigates visual analogical reasoning in large multimodal models (LMMs) compared to human adults and children. A "visual analogy" is an abstract rule inferred from one image and applied to another. While benchmarks exist for testing visual reasoning in LMMs, they require advanced skills and omit basic visual analogies that even young children can make. Inspired by developmental psychology, we propose a new benchmark of 4,300 visual transformations of everyday objects to test LMMs on visual analogical reasoning and compare them to children (ages three to five) and to adults. We structure the evaluation into three stages: identifying what changed (e.g., color, number, etc.), how it changed (e.g., added one object), and applying the rule to new scenarios. Our findings show that while GPT-o1, GPT-4V, LLaVA-1.5, and MANTIS identify the "what" effectively, they struggle with quantifying the "how" and extrapolating this rule to new objects. In contrast, children and adults exhibit much stronger analogical reasoning at all three stages. Additionally, the strongest tested model, GPT-o1, performs better in tasks involving simple surface-level visual attributes like color and size, correlating with quicker human adult response times. Conversely, more complex tasks such as number, rotation, and reflection, which necessitate extensive cognitive processing and understanding of extrinsic spatial properties in the physical world, present more significant challenges. Altogether, these findings highlight the limitations of training models on data that primarily consists of 2D images and text. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9245a6ca-1008-47ae-859f-a99e165e5423Cited by top-tier papers7
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata et al.NeurIPS 2024 · 101 citations
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual SimulationsLinjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma et al.ICLR 2026 · 22 citations
- Visual symbolic mechanisms: Emergent symbol processing in Vision Language ModelsRim Assouel, Declan Iain Campbell, Yoshua Bengio, Taylor Whittington WebbICLR 2026 · 14 citations
- PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in PuzzlehuntsHengzhi Li, Justin Zhang, Brendon Jiang, Alexander Naehu et al.ICLR 2026 · 6 citations
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation ModelsShengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton et al.CVPR 2026 · 4 citations
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao et al.NeurIPS 2023 · 810 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson et al.NeurIPS 2022 · 340 citations
Related papers
- VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical ReasoningNilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale et al.ICLR 2025
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMsRohit Sinha, Aditya Sanjiv Kanade, Sai Srinivas Kancheti, Vineeth N. Balasubramanian et al.ACL 2026
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeYueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li et al.ICML 2026 · 44 citations
- Easy for Children, Hard for AI: The Limits of Multimodal LLMs in Early Childhood LearningJingping Liu, Xueyan Wu, Hanxuan Chen, Ziyan Liu et al.AAAI 2026
