SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
Xinyi He, Qian Liu, Mingzhe Du, Lin Yan, ZhiJie Fan, Yiming Huang, Yin Zheng, Zejian Yuan, Zejun MA
Abstract
Code performance optimization is paramount in real-world software engineering and critical for production-level systems. While Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and bug fixing, their proficiency in enhancing code performance at the repository level remains largely unexplored. To address this gap, we introduce SWE-Perf, the first benchmark specifically designed to systematically evaluate LLMs on code performance optimization tasks within authentic repository contexts. SWE-Perf comprises 140 carefully curated instances, each derived from performance-improving pull requests from popular GitHub repositories. Each benchmark instance includes the relevant codebase, target functions, performance-related tests, expert-authored patches, and executable environments. Through a comprehensive evaluation of representative methods that span file-level and repo-level approaches (e.g., Agentless and Open-Hands), we reveal a substantial capability gap between existing LLMs and expert-level optimization performance, highlighting critical research opportunities in this emerging field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationMingzhe Du, Anh Tuan Luu, Yue Liu, Yuhao Qing et al.NeurIPS 2025 · 18 citations
- EvoClaw: Evaluating AI Agents on Continuous Software EvolutionGangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan et al.ICML 2026 · 6 citations
- CodeClash: Benchmarking Goal-Oriented Software EngineeringJohn Yang, Kilian Lieret, Joyce Yang, Carlos Jimenez et al.ICML 2026 · 5 citations
- QuArch: A Benchmark for Evaluating LLM Reasoning in Computer ArchitectureShvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand et al.ICML 2026 · 3 citations
- MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software EngineeringChuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen et al.ICML 2026 · 3 citations
Builds on9
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon et al.ICLR 2024 · 141 citations
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu et al.ICLR 2025 · 7 citations
Related papers
- FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature ImplementationWei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao et al.ACL 2025 · 40 citations
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
- SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?Jeffrey Ma, Milad Hashemi, Amir Yazdanbakhsh, Kevin Swersky et al.ICML 2026 · 13 citations
- FormulaCode: Evaluating Agentic Optimization on Large CodebasesAtharva Sehgal, James Hou, Akanksha Sarkar, Ishaan Mantripragada et al.ICML 2026 · 3 citations
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu et al.ICML 2026 · 9 citations
