SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
Jingxuan Xu, Ken Deng, Weihao Li, Songwei Yu, Haoyang Huang, Xinping Lei, Yifan Yao, Huaixi Tang, Zhiyi Lai, Kepeng Lei, Zizheng Zhan, Yanan Wu
Abstract
Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Pythoncentric bug fixing, leaving critical dimensions of software engineering underexplored. To address these gaps, we introduce SWE-Compass 1 , a comprehensive benchmark that unifies heterogeneous code-related evaluations into a structured and productionaligned framework. SWE-Compass spans 8 task types, 8 programming scenarios, and 10 programming languages, with 2000 high-quality instances curated from authentic GitHub pull requests and refined through systematic filtering and validation. We benchmark ten state-of-the-art LLMs under two agentic frameworks, SWE-Agent and Claude Code, revealing a clear hierarchy of difficulty across task types, languages, and scenarios. Moreover, by aligning evaluation with real-world developer practices, we hope SWE-Compass can provide a rigorous and reproducible foundation for diagnosing and advancing agentic coding capabilities in large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63c4a7a1-1207-4c25-82ae-68843e2de4aaCited by top-tier papers3
- CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative TournamentsLingyue Fu, Xin Ding, Linyue Pan, Yaoming Zhu et al.ICML 2026 · 3 citations
- CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level TasksLingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan et al.ACL 2026
- AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering AgentsZhengran Zeng, Yixin Li, Rui Xie, Wei Ye et al.ISSTA 2026
Builds on12
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang et al.ICLR 2023 · 140 citations
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
Related papers
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao et al.ICLR 2026 · 30 citations
- Unified Software Engineering Agent as AI Software EngineerLeonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang et al.ICSE 2026
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan et al.ICML 2026 · 31 citations
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong et al.ASE 2024 · 7 citations
