CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
Natasha Butt, Blazej Manczak, Auke J. Wiggers, Corrado Rainone, David W. Zhang, Michaël Defferrard, Taco Cohen
Abstract
Large language models are increasingly solving tasks that are commonly believed to require human-level reasoning ability. However, these models still perform very poorly on benchmarks of general intelligence such as the Abstraction and Reasoning Corpus (ARC). In this paper, we approach ARC as a programming-by-examples problem, and introduce a novel and scalable method for language model self-improvement called Code Iteration (CodeIt). Our method iterates between 1) program sampling and hindsight relabeling, and 2) learning from prioritized experience replay. By relabeling the goal of an episode (i.e., the target program output given input) to the realized output produced by the sampled program, our method effectively deals with the extreme sparsity of rewards in program synthesis. Applying CodeIt to the ARC dataset, we demonstrate that prioritized hindsight replay, along with pre-training and data-augmentation, leads to successful inter-task generalization. CodeIt is the first neuro-symbolic approach that scales to the full ARC evaluation dataset. Our method solves 15% of ARC evaluation tasks, achieving state-of-the-art performance and outperforming existing neural and symbolic baselines. Our code is available at https://github.com/ Qualcomm-AI-research/codeit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c948981-bef0-4bdf-9675-f9e1da0c904dCited by top-tier papers9
- Searching Latent Program SpacesMatthew Macfarlane, Clément BonnetNeurIPS 2025 · 23 citations
- SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesDongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir et al.AAAI 2025 · 8 citations
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao et al.CVPR 2026
- Combining Induction and Transduction for Abstract ReasoningWen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu et al.ICLR 2025
- Self-Training Large Language Models for Improved Visual Program Synthesis With Visual ReinforcementZaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu et al.CVPR 2024
Builds on11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang et al.ICML 2024 · 443 citations
- Hypothesis Search: Inductive Reasoning with Language ModelsRuocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu et al.ICLR 2024 · 156 citations
- Learning to Prove Theorems by Learning to Generate TheoremsMingzhe Wang, Jia DengNeurIPS 2020 · 60 citations
Related papers
- Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of PerspectiveDaniel Franzen, Jan Disselhoff, David HartmannICML 2025
- Generalized Planning for the Abstraction and Reasoning CorpusChao Lei, Nir Lipovetzky, Krista A. EhingerAAAI 2024 · 13 citations
- ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)Kartik Singhal, Gautam ShroffAAAI 2025
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
- ANPL: Towards Natural Programming with Interactive DecompositionDi Huang, Ziyuan Nan, Xing Hu, Pengwei Jin et al.NeurIPS 2023 · 24 citations
