Language Models Can Teach Themselves to Program Better
Patrick Haluptzok, Matthew Bowers, Adam Tauman Kalai
Abstract
Recent Language Models (LMs) achieve breakthrough performance in code generation when trained on human-authored problems, even solving some competitive-programming problems. Self-play has proven useful in games such as Go, and thus it is natural to ask whether LMs can generate their own instructive programming problems to improve their performance. We show that it is possible for an LM to synthesize programming problems and solutions, which are filtered for correctness by a Python interpreter. The LM's performance is then seen to improve when it is fine-tuned on its own synthetic problems and verified solutions; thus the model "improves itself" using the Python interpreter. Problems are specified formally as programming puzzles [Schuster et al., 2021] , a code-based problem format where solutions can easily be verified for correctness by execution. In experiments on publicly-available LMs, test accuracy more than doubles. This work demonstrates the potential for code LMs, with an interpreter, to generate instructive problems and improve their own performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2493f032-e14e-445b-84e8-d9a69769f163Cited by top-tier papers39
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsLing Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao et al.NeurIPS 2024 · 144 citations
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon et al.ICLR 2024 · 141 citations
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar et al.ICLR 2024 · 114 citations
Builds on6
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- DreamCoder: bootstrapping inductive program synthesis with wake-sleep library learningKevin Ellis, Catherine Wong, Maxwell I. Nye, Mathias Sablé-Meyer et al.PLDI 2021 · 97 citations
- LIME: Learning Inductive Bias for Primitives of Mathematical ReasoningYuhuai Wu, Markus N. Rabe, Wenda Li, Jimmy Ba et al.ICML 2021 · 66 citations
Related papers
- VERSE: Verification-based Self-Play for Code InstructionsHao Jiang, Qi Liu, Rui Li, Yuze Zhao et al.AAAI 2025 · 3 citations
- Propose, Solve, Verify: Self-Play Through Formal VerificationAlex Wilf, Pranjal Aggarwal, Bryan Parno, Daniel Fried et al.ICML 2026 · 6 citations
- ACES: Generating a Diversity of Challenging Programming Puzzles with Autotelic Generative ModelsJulien Pourcel, Cédric Colas, Gaia Molinaro, Pierre-Yves Oudeyer et al.NeurIPS 2024 · 9 citations
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen et al.ICLR 2026 · 12 citations
- Revisit Self-Debugging with Self-Generated Tests for Code GenerationXiancai Chen, Zhengwei Tao, Kechi Zhang, Changzhi Zhou et al.ACL 2025
