CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
Natasha Butt, Blazej Manczak, Auke J. Wiggers, Corrado Rainone, David W. Zhang, Michaël Defferrard, Taco Cohen
摘要
Large language models are increasingly solving tasks that are commonly believed to require human-level reasoning ability. However, these models still perform very poorly on benchmarks of general intelligence such as the Abstraction and Reasoning Corpus (ARC). In this paper, we approach ARC as a programming-by-examples problem, and introduce a novel and scalable method for language model self-improvement called Code Iteration (CodeIt). Our method iterates between 1) program sampling and hindsight relabeling, and 2) learning from prioritized experience replay. By relabeling the goal of an episode (i.e., the target program output given input) to the realized output produced by the sampled program, our method effectively deals with the extreme sparsity of rewards in program synthesis. Applying CodeIt to the ARC dataset, we demonstrate that prioritized hindsight replay, along with pre-training and data-augmentation, leads to successful inter-task generalization. CodeIt is the first neuro-symbolic approach that scales to the full ARC evaluation dataset. Our method solves 15% of ARC evaluation tasks, achieving state-of-the-art performance and outperforming existing neural and symbolic baselines. Our code is available at https://github.com/ Qualcomm-AI-research/codeit .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Searching Latent Program SpacesMatthew Macfarlane, Clément BonnetNeurIPS 2025 · 被引用 23 次
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesDongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir 等AAAI 2025 · 被引用 8 次
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao 等CVPR 2026
- Combining Induction and Transduction for Abstract ReasoningWen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu 等ICLR 2025
- Self-Training Large Language Models for Improved Visual Program Synthesis With Visual ReinforcementZaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu 等CVPR 2024
它引用的顶会 Paper11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang 等ICML 2024 · 被引用 443 次
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu 等ICLR 2024 · 被引用 156 次
- Learning to Prove Theorems by Learning to Generate TheoremsMingzhe Wang, Jia DengNeurIPS 2020 · 被引用 60 次
相关 Paper
- Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of PerspectiveDaniel Franzen, Jan Disselhoff, David HartmannICML 2025
- Generalized Planning for the Abstraction and Reasoning CorpusChao Lei, Nir Lipovetzky, Krista A. EhingerAAAI 2024 · 被引用 13 次
- ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)Kartik Singhal, Gautam ShroffAAAI 2025
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella 等ICML 2025
- ANPL: Towards Natural Programming with Interactive DecompositionDi Huang, Ziyuan Nan, Xing Hu, Pengwei Jin 等NeurIPS 2023 · 被引用 24 次
