Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
Jenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange, Jeff Clune
Abstract
Most of today's AI systems are constrained by human-designed, fixed architectures and cannot autonomously and continuously improve themselves. The scientific method, on the other hand, is a cumulative and open-ended system, where each innovation builds upon previous artifacts, enabling future discoveries. There is growing hope that the current manual process of advancing AI could itself be automated. If done safely, such automation would accelerate AI development and allow us to reap its benefits much sooner. This prospect raises the question of how AI systems can endlessly improve themselves while getting better at solving relevant problems. Meta-learning can automate the discovery of novel algorithms, but is limited by first-order improvements and the human design of a suitable search space. The Gödel machine (Schmidhuber, 2007) proposed a theoretical alternative: a self-improving AI that repeatedly modifies itself in a provably beneficial manner. Unfortunately, proving that most changes are net beneficial is impossible in practice. We introduce the Darwin Gödel Machine (DGM), a novel self-improving system that iteratively modifies its own code (thereby also improving its ability to modify its own codebase) and empirically validates each change using coding benchmarks. Inspired by Darwinian evolution and open-endedness research, the DGM grows an archive of generated coding agents. It samples agents from this archive, which self-modify to create new, interesting versions of themselves. This open-ended exploration forms a growing tree of diverse, high-quality agents and allows the parallel exploration of many different paths through the search space. Empirically, the DGM automatically improves its coding capabilities (e.g., better code editing tools, long-context window management, peer-review mechanisms), increasing performance on SWE-bench from 20.0% to 50.0%, and on Polyglot from 14.2% to 30.7%. Furthermore, the DGM significantly outperforms baselines without selfimprovement or open-ended exploration. All experiments were done with safety precautions (e.g., sandboxing, human oversight). Overall, the DGM represents a significant step toward self-improving AI, capable of gathering its own stepping stones along a path that unfolds into endless innovation. All code is open-sourced at https://github.com/jennyzzt/dgm .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 301212ad-56f5-4d6b-ad62-7558bda25378Cited by top-tier papers6
- VeRO: A Harness for Agents to Optimize AgentsVarun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue et al.ICML 2026 · 6 citations
- Procedural Generation Of Algorithm Discovery Tasks in Machine LearningAlexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani et al.ICML 2026 · 2 citations
- Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-ImprovementXinmeng Hou, Bohao Qu, Wuqi Wang, Peiliang Gong et al.ACL 2026 · 1 citation
- Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel GenerationYee Hin Chong, Jiaming Wu, Youhui Zhang, Peng QuICML 2026
- Towards Scalable Oversight via Partitioned Human SupervisionRen Yin, Takashi Ishida, Masashi SugiyamaICLR 2026
Builds on48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
Related papers
- Huxley-Göodel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving MachineWenyi Wang, Piotr Piekos, Li Nanbo, Firas Laakom et al.ICLR 2026 · 15 citations
- Gödel Agent: A Self-Referential Agent Framework for Recursively Self-ImprovementXunjian Yin, Xinyi Wang, Liangming Pan, Li Lin et al.ACL 2025 · 7 citations
- Automated Design of Agentic SystemsShengran Hu, Cong Lu, Jeff CluneICLR 2025
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han et al.NeurIPS 2025 · 73 citations
- Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent EvolutionHongze Mi, Yibo Feng, Wenjie Lu, Song Cao et al.ICML 2026
