ARC Is a Vision Problem!
Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding, Runqian Wang, Yeyin Eva Zhu, Jacob Andreas, Kaiming He
Abstract
The Abstraction and Reasoning Corpus (ARC) is designed to promote research on abstract reasoning, a fundamental aspect of human intelligence. Common approaches to ARC treat it as a language-oriented problem, addressed by large language models (LLMs) or recurrent reasoning models. However, although the puzzle-like tasks in ARC are inherently visual, existing research has rarely approached the problem from a vision-centric perspective. In this work, we formulate ARC within a vision paradigm, framing it as an image-to-image translation problem. To incorporate visual priors, we represent the inputs on a “canvas” that can be processed like natural images.It is then straightforward for us to apply standard vision architectures, such as a vanilla Vision Transformer (ViT), to perform image-to-image mapping. Our model is trained from scratch solely on ARC data and generalizes to unseen tasks through test-time training. Our framework, termed Vision ARC (VARC), achieves 60.4% accuracy on the ARC-1 benchmark, substantially outperforming existing methods that are also trained from scratch. Our results are competitive with those of leading LLMs and close the gap to average human performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01deea07-3ef5-4503-9c51-c1ba9d0fbd9eCited by top-tier papers2
- Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning EfficiencyMingqing Xiao, Yansen Wang, Dongqi Han, Caihua Shan et al.ICML 2026 · 2 citations
- Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic ReasoningBowen LIU, Zhi Wu, RunquanXie, Zhanhui Kang et al.ICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao et al.CVPR 2026
- Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning BenchmarksXinhe Wang, Jin Huang, Xingjian Zhang, Tianhao Wang et al.ACL 2026 · 3 citations
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu et al.ICLR 2024 · 156 citations
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language ModelsJiayun Luo, Wan-Cyuan Fan, Lyuyang Wang, Xiangteng He et al.ICLR 2026 · 19 citations
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery SimulationPhillip Y. Lee, Jihyeon Je, Chanho Park, Leonidas J. Guibas et al.ICCV 2025 · 6 citations
