Huxley-Göodel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
Wenyi Wang, Piotr Piekos, Li Nanbo, Firas Laakom, Yimeng Chen, Mateusz Ostaszewski, Mingchen Zhuge, Jürgen Schmidhuber
Abstract
Recent studies operationalize self-improvement through coding agents that edit their own codebases. They grow a tree of self-modifications through expansion strategies that favor higher software engineering benchmark performance, assuming that this implies more promising subsequent self-modifications. However, we identify a mismatch between the agent’s self-improvement potential (metaproductivity) and its coding benchmark performance, namely the Metaproductivity-Performance Mismatch. Inspired by Huxley’s concept of clade, we propose a metric () that aggregates the benchmark performances of the descendants of an agent as an indicator of its potential for self-improvement. We show that, in our self-improving coding agent development setting, access to the true CMP is sufficient to simulate how the Gödel Machine would behave under certain assumptions. We introduce the Huxley-Gödel Machine (HGM), which, by estimating and using it as guidance, searches the tree of self-modifications. On SWE-bench Verified and Polyglot, HGM outperforms prior self-improving coding agent development methods while using fewer allocated CPU hours. Last but not least, HGM demonstrates strong transfer to other coding datasets and LLMs. %large language models. The agent optimized by HGM on SWE-bench Verified with GPT-5 mini and evaluated on SWE-bench Lite with GPT-5 achieves human-level performance, matching the best officially checked results of human-engineered coding agents. Our code is publicly available at https://github.com/metauto-ai/HGM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c086b68c-1312-4f3e-a1da-0ba593cdc6a4Cited by top-tier papers2
- Procedural Generation Of Algorithm Discovery Tasks in Machine LearningAlexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani et al.ICML 2026 · 2 citations
- A Task-centric Theory for Iterative Self-Improvement with Easy-to-Hard CurriculaChenruo Liu, Yijun Dong, Yiqiu Shen, Qi LeiICML 2026
Builds on6
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
- Demystifying LLM-Based Software Engineering AgentsChunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming ZhangFSE 2025 · 36 citations
- Discovering Evolution Strategies via Meta-Black-Box OptimizationRobert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy et al.ICLR 2023 · 21 citations
Related papers
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange et al.ICLR 2026 · 101 citations
- Gödel Agent: A Self-Referential Agent Framework for Recursively Self-ImprovementXunjian Yin, Xinyi Wang, Liangming Pan, Li Lin et al.ACL 2025 · 7 citations
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao et al.ICLR 2026 · 30 citations
- Kimi-Dev: Agentless Training as Skill Prior for SWE-agentsZonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He et al.ICLR 2026 · 34 citations
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan et al.ICML 2026 · 31 citations
