Unified Software Engineering Agent as AI Software Engineer
Leonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang, Lin Tan, Abhik Roychoudhury
摘要
The growth of Large Language Model (LLM) technology has raised expectations for automated coding. However, software engineering is more than coding and is concerned with activities including maintenance and evolution of a project. In this context, the concept of LLM agents has gained traction, which utilize LLMs as reasoning engines to invoke external tools autonomously. But is an LLM agent the same as an AI software engineer? In this paper, we seek to understand this question by developing a Unified Software Engineering agent or USEagent. Unlike existing work which builds specialized agents for specific software tasks such as testing, debugging, and repair, our goal is to build a unified agent which can orchestrate and handle multiple capabilities. This gives the agent the promise of handling complex scenarios in software development such as fixing an incomplete patch, adding new features, or taking over code written by others. We envision USEagent as the first draft of a future AI Software Engineer which can be a team member in future software development teams involving both AI and humans. To evaluate the efficacy of USEagent, we build a Unified Software Engineering bench (USEbench) comprising of myriad tasks such as coding, testing, and patching. USEbench is a judicious mixture of tasks from existing benchmarks such as SWE-bench, SWT-bench, and REPOCOD. In an evaluation on USEbench consisting of 1,271 repository-level software engineering tasks, USEagent shows improved efficacy compared to existing general agents such as OpenHands CodeActAgent. For specific tasks such as issue resolution in SWE-bench, USEagent demonstrates efficacy comparable to agents such as AutoCodeRover (which are focused on software maintenance), while maintaining applicability to a broader range of software engineering tasks. There exist gaps in the capabilities of USEagent for certain coding tasks, which provides hints on further developing the AI Software Engineer of the future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Structurally Aligned Subtask-Level Memory for Software Engineering AgentsKangning Shen, Jingyuan Zhang, Chenxi Sun, Wencong Zeng 等ICML 2026 · 被引用 9 次
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu 等ICML 2026 · 被引用 4 次
- Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and SelectionToufique Ahmed, Jatin Ganhotra, Avraham Shinnar, Martin HirzelICSE 2026 · 被引用 2 次
- AutoCodeSherpa: Symbolic Explanations in AI Coding AgentsSungmin Kang, Haifeng Ruan, Abhik RoychoudhuryISSTA 2026
- Can Old Tests Do New Tricks for Resolving SWE Issues?Yang Chen, Toufique Ahmed, Reyhaneh Jabbarvand, Martin HirzelFSE 2026
它引用的顶会 Paper14
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang 等ICML 2024 · 被引用 436 次
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
相关 Paper
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen 等NeurIPS 2025 · 被引用 4 次
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 被引用 96 次
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu 等ICML 2026 · 被引用 9 次
- RepoGraph: Enhancing AI Software Engineering with Repository-level Code GraphSiru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao 等ICLR 2025
- CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding ChallengesKechi Zhang, Jia Li, Ge Li, Xianjie Shi 等ACL 2024
