Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
Daniel Franzen, Jan Disselhoff, David Hartmann
摘要
The Abstraction and Reasoning Corpus (ARC-AGI) poses a significant challenge for large language models (LLMs), exposing limitations in their abstract reasoning abilities. In this work, we leverage task-specific data augmentations throughout the training, generation, and scoring phases, and employ a depth-first search algorithm to generate diverse, high-probability candidate solutions. Furthermore, we utilize the LLM not only as a generator but also as a scorer, using its output probabilities to select the most promising solutions. Our method achieves a score of 71.6% (286.5/400 solved tasks) on the public ARC-AGI evaluation set, demonstrating state-of-the-art performance among publicly available approaches. While concurrent closed-source work has reported higher scores, our method distinguishes itself through its transparency, reproducibility, and remarkably low inference cost, averaging only around 2ct per task on readily available hardware. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ARC Is a Vision Problem!Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding 等CVPR 2026 · 被引用 23 次
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao 等CVPR 2026
- One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion ModelsChris Cameron, Wangzheng Wang, Nikita Ivanov, Ashmita Bhattacharyya 等ICML 2026
- Compositional Generalization through Gradient Search in Nonparametric Latent SpaceHaruki Shirakami, James HendersonICLR 2026
- Context Tuning for In-Context OptimizationJack Lu, Ryan Teehan, Zhenbang Yang, Mengye RenICML 2026
它引用的顶会 Paper4
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- Physics of Language Models: Part 3.2, Knowledge ManipulationZeyuan Allen-Zhu, Yuanzhi LiICLR 2025 · 被引用 2 次
- Combining Induction and Transduction for Abstract ReasoningWen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu 等ICLR 2025
相关 Paper
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu 等ICLR 2024 · 被引用 156 次
- CodeIt: Self-Improving Language Models with Prioritized Hindsight ReplayNatasha Butt, Blazej Manczak, Auke J. Wiggers, Corrado Rainone 等ICML 2024 · 被引用 29 次
- ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)Kartik Singhal, Gautam ShroffAAAI 2025
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu 等ICML 2025
- SatLM: Satisfiability-Aided Language Models Using Declarative PromptingXi Ye, Qiaochu Chen, Isil Dillig, Greg DurrettNeurIPS 2023 · 被引用 126 次
