A large-scale benchmark for few-shot program induction and synthesis
Ferran Alet, Javier Lopez-Contreras, James Koppel, Maxwell I. Nye, Armando Solar-Lezama, Tomás Lozano-Pérez, Leslie Pack Kaelbling, Joshua B. Tenenbaum
Abstract
A landmark challenge for AI is to learn flexible, powerful representations from small numbers of examples. On an important class of tasks, hypotheses in the form of programs provide extreme generalization capabilities from surprisingly few examples. However, whereas large real image benchmarks have spurred progress in metalearning for deep networks, there is no comparably big, real program-synthesis dataset. This is because, while images are relatively easy to label from internet meta-data or annotated by nonexperts, generating meaningful input-output tests for program induction has proven hard to scale. In this work, we propose a new way of leveraging a collection of programs with associated unit tests to create a much larger collection of testprogram pairs. We do so by extracting subprograms of each program and using the inputs of the overall program to get tests for each subprogram. This allows us to create PROGRES, a large-scale few-shot program-induction benchmark of real programs and propose new challenges in this domain. We analyze the effect of multiple design choices on transformer-based program induction and synthesis algorithms, pointing to shortcomings of current methods and suggesting multiple avenues for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar et al.ICLR 2024 · 114 citations
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
- Latent Execution for Neural Program Synthesis Beyond Domain-Specific LanguagesXinyun Chen, Dawn Song, Yuandong TianNeurIPS 2021 · 56 citations
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu et al.ICLR 2026 · 35 citations
- Language Models Can Teach Themselves to Program BetterPatrick Haluptzok, Matthew Bowers, Adam Tauman KalaiICLR 2023 · 17 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
Related papers
- ExeDec: Execution Decomposition for Compositional Generalization in Neural Program SynthesisKensen Shi, Joey Hong, Yinlin Deng, Pengcheng Yin et al.ICLR 2024 · 21 citations
- Is Programming by Example Solved by LLMs?Wen-Ding Li, Kevin EllisNeurIPS 2024 · 45 citations
- Think Big, Teach Small: Do Language Models Distil Occam's Razor?Gonzalo Jaimovitch-López, David Castellano Falcón, César Ferri, José Hernández-OralloNeurIPS 2021 · 3 citations
- Learning to Prove Theorems by Learning to Generate TheoremsMingzhe Wang, Jia DengNeurIPS 2020 · 60 citations
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim et al.NeurIPS 2025 · 4 citations
