MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Aleksander Madry, Lilian Weng
摘要
This paper develops a theory of search stability for long-running agents operating under finite active context, delayed verification, sparse expensive feedback, path-dependent lock-in, and lossy state compression. The focus is not only on model quality, but on the mesoscopic law layer that governs how an agent should preserve, retire, substitute, compress, branch, and reset competing hypotheses or route summaries over time. The framework models search state as an active hypothesis portfolio partitioned into coarse families under a context budget. Each item carries promise, verification lag, retention cost, staleness, overlap burden, and inertia. A central contribution is a set-valued adequacy semantics: within each discrimination window, the system is associated with a nonempty random set of operationally adequate families induced by the realized initial information state and downstream randomness. Success is defined as preserving recoverability of at least one adequate family at the first strongly discriminating verification stage, avoiding dependence on a selector-defined pseudo-truth. The paper derives threshold and impossibility results for context contamination, shadow retirement, delayed-verification coverage, reserve feasibility, and budget-limited adequacy. It also develops a theory of within-family semantic substitution, compressed-control alias hazard, reset admissibility, stale-legacy drift, diagnostic regret decomposition, and rolling-window lifting for long-running agents with repeated verification stages and changing task modes. The intended contribution is an audit-and-design law layer for bounded-memory AI systems. The theory is deliberately narrow and conditional, but it aims to make long-horizon agent failures more diagnosable: separating failures caused by bounded-memory hypothesis ecology from failures caused by raw model weakness, and from mixtures of both.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng 等NeurIPS 2025 · 被引用 160 次
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 被引用 86 次
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra 等NeurIPS 2025 · 被引用 71 次
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree SearchYuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki 等NeurIPS 2025 · 被引用 67 次
它引用的顶会 Paper7
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 被引用 209 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
相关 Paper
- Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning ModelsSho Sonoda, Shunta Akiyama, Yuya UezatoICML 2026 · 被引用 2 次
- Deep Search with Hierarchical Meta-Cognitive Monitoring Inspired by Cognitive NeuroscienceZhongxiang Sun, Qipeng Wang, Weijie Yu, Jingxuan Yang 等SIGIR 2026
- Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-AuditingWenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu 等ACL 2026 · 被引用 3 次
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 被引用 246 次
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
