Abstract Visual Reasoning with Tangram Shapes
Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert D. Hawkins, Yoav Artzi
Abstract
We introduce KiloGram, a resource for studying abstract visual reasoning in humans and machines. Drawing on the history of tangram puzzles as stimuli in cognitive science, we build a richly annotated dataset that, with >1k distinct stimuli, is orders of magnitude larger and more diverse than prior resources. It is both visually and linguistically richer, moving beyond whole shape descriptions to include segmentation maps and part labels. We use this resource to evaluate the abstract visual reasoning capacities of recent multi-modal models. We observe that pre-trained weights demonstrate limited abstract reasoning, which dramatically improves with fine-tuning. We also observe that explicitly describing parts aids abstract reasoning for both humans and models, especially when jointly encoding the linguistic and visual inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Large Language Models are Visual Reasoning CoordinatorsLiangyu Chen, Bo Li, Sheng Shen, Jingkang Yang et al.NeurIPS 2023 · 108 citations
- CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsZihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei et al.AAAI 2025 · 36 citations
- MEWL: Few-shot multimodal word learning with referential uncertaintyGuangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang et al.ICML 2023 · 29 citations
- GThinker: Towards General Multimodal Reasoning via Cue-Guided RethinkingYufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue et al.CVPR 2026 · 14 citations
Builds on3
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
Related papers
- Learning from the Tangram to Solve Mini Visual TasksYizhou Zhao, Liang Qiu, Pan Lu, Feng Shi et al.AAAI 2022 · 5 citations
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- PTR: A Benchmark for Part-based Conceptual, Relational, and Physical ReasoningYining Hong, Li Yi, Josh Tenenbaum, Antonio Torralba et al.NeurIPS 2021 · 46 citations
- PARTONOMY: Large Multimodal Models with Part-Level Visual UnderstandingAnsel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan et al.NeurIPS 2025 · 5 citations
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 35 citations
