Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoning
Beichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang
Abstract
reasoning from minimal examples remains a core unsolved problem for frontier foundation models such as GPT-5. These models still fail to infer structured transformation rules from a handful of examples, which is a key hallmark of human intelligence. The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) provides a rigorous testbed for this capability, demanding conceptual rule induction and transfer to novel tasks. Most existing methods treat ARC-AGI as a purely textual reasoning task, overlooking the fact that humans rely heavily on visual abstraction when solving such puzzles. However, our experiments reveal a paradox: naively rendering ARC-AGI grids as images degrades performance due to imprecise rule execution. This leads to our central hypothesis that vision and language possess complementary strengths across distinct reasoning stages: vision supports global pattern abstraction and verification, whereas language specializes in symbolic rule formulation and precise execution. Building on this insight, we introduce two synergistic strategies: (1) Vision-Language Synergy Reasoning (VLSR), which decomposes ARC-AGI into modality-aligned subtasks; and (2) Modality-Switch Self-Correction (MSSC), which leverages vision to verify text-based reasoning for intrinsic error correction. Extensive experiments demonstrate that our approach yields up to a 4.33% improvement over text-only baselines across diverse flagship models and multiple ARC-AGI tasks. Our findings suggest that unifying visual abstraction with linguistic reasoning is a crucial step toward achieving generalizable, human-like intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58f91a24-6363-4feb-814b-c19eb1aef4fdBuilds on15
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu et al.ICLR 2024 · 156 citations
Related papers
- ARC Is a Vision Problem!Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding et al.CVPR 2026 · 23 citations
- Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning BenchmarksXinhe Wang, Jin Huang, Xingjian Zhang, Tianhao Wang et al.ACL 2026 · 3 citations
- Human-like Abstract Visual Reasoning via Understanding and Solving Reasoning LoopXinwang Chen, Xiuxing Li, Qing Li, Ziyue Zhuang et al.CVPR 2026
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji et al.ICCV 2025 · 1 citation
- Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu et al.ICML 2025
