Is Programming by Example Solved by LLMs?
Wen-Ding Li, Kevin Ellis
摘要
Programming-by-Examples (PBE) aims to generate an algorithm from input-output examples. Such systems are practically and theoretically important: from an end-user perspective, they are deployed to millions of people, and from an AI perspective, PBE corresponds to a very general form of few-shot inductive inference. Given the success of Large Language Models (LLMs) in code-generation tasks, we investigate here the extent to which LLMs can be said to have"solved"PBE. We experiment on classic domains such as lists and strings, and an uncommon graphics programming domain not well represented in typical pretraining data. We find that pretrained models are not effective at PBE, but that they can be fine-tuned for much higher performance, provided the test problems are in-distribution. We analyze empirically what causes these models to succeed and fail, and take steps toward understanding how to achieve better out-of-distribution generalization. Collectively these results suggest that LLMs make strong progress toward solving the typical suite of PBE tasks, potentially increasing the flexibility and applicability of PBE systems, while also identifying ways in which LLMs still fall short.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Tikzero: Zero-Shot Text-Guided Graphics Program SynthesisJonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka 等ICCV 2025 · 被引用 24 次
- Generating Computational Cognitive models using Large Language ModelsMilena Rmus, Akshay Kumar Jagadish, Marvin Mathony, Tobias Ludwig 等NeurIPS 2025 · 被引用 19 次
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang 等ACL 2026 · 被引用 5 次
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim 等NeurIPS 2025 · 被引用 4 次
- The First Prompt Counts the Most! An Evaluation of Large Language Models on Iterative Example-Based Code GenerationYingjie Fu, Bozhou Li, Linyi Li, Wentao Zhang 等ISSTA 2025 · 被引用 3 次
它引用的顶会 Paper28
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- PRM-PBE: Process Reward Model for Reinforcement Learning in Programming-by-ExampleYue Fang, Zhi Jin, Jie An, Hongshen Chen 等ICML 2026
- Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python CodeAugusto B. Corrêa, André Grahl Pereira, Jendrik SeippNeurIPS 2025 · 被引用 27 次
- Programming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law InductionAtharva Naik, Darsh Agrawal, Hong Sng, Clayton Marr 等ACL 2025 · 被引用 1 次
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang 等EMNLP 2022 · 被引用 103 次
- Think Big, Teach Small: Do Language Models Distil Occam's Razor?Gonzalo Jaimovitch-López, David Castellano Falcón, César Ferri, José Hernández-OralloNeurIPS 2021 · 被引用 3 次
