Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
Daniel Franzen, Jan Disselhoff, David Hartmann
Abstract
The Abstraction and Reasoning Corpus (ARC-AGI) poses a significant challenge for large language models (LLMs), exposing limitations in their abstract reasoning abilities. In this work, we leverage task-specific data augmentations throughout the training, generation, and scoring phases, and employ a depth-first search algorithm to generate diverse, high-probability candidate solutions. Furthermore, we utilize the LLM not only as a generator but also as a scorer, using its output probabilities to select the most promising solutions. Our method achieves a score of 71.6% (286.5/400 solved tasks) on the public ARC-AGI evaluation set, demonstrating state-of-the-art performance among publicly available approaches. While concurrent closed-source work has reported higher scores, our method distinguishes itself through its transparency, reproducibility, and remarkably low inference cost, averaging only around 2ct per task on readily available hardware. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 834d55f8-4052-45ce-8b48-24449cf1521dCited by top-tier papers5
- ARC Is a Vision Problem!Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding et al.CVPR 2026 · 23 citations
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao et al.CVPR 2026
- One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion ModelsChris Cameron, Wangzheng Wang, Nikita Ivanov, Ashmita Bhattacharyya et al.ICML 2026
- Compositional Generalization through Gradient Search in Nonparametric Latent SpaceHaruki Shirakami, James HendersonICLR 2026
- Context Tuning for In-Context OptimizationJack Lu, Ryan Teehan, Zhenbang Yang, Mengye RenICML 2026
Builds on4
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 258 citations
- Physics of Language Models: Part 3.2, Knowledge ManipulationZeyuan Allen-Zhu, Yuanzhi LiICLR 2025 · 2 citations
- Combining Induction and Transduction for Abstract ReasoningWen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu et al.ICLR 2025
Related papers
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu et al.ICLR 2024 · 156 citations
- CodeIt: Self-Improving Language Models with Prioritized Hindsight ReplayNatasha Butt, Blazej Manczak, Auke J. Wiggers, Corrado Rainone et al.ICML 2024 · 29 citations
- ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)Kartik Singhal, Gautam ShroffAAAI 2025
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu et al.ICML 2025
- SatLM: Satisfiability-Aided Language Models Using Declarative PromptingXi Ye, Qiaochu Chen, Isil Dillig, Greg DurrettNeurIPS 2023 · 126 citations
