SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes Heidecke
Abstract
We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at 50 bug fixes to $32,000 feature implementations -and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with endto-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond ( https://github.com/ openai/SWELancer-Benchmark ). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c78c709c-c791-4ae4-88f6-733d9bff2f8bCited by top-tier papers18
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li et al.ICLR 2026 · 520 citations
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng et al.NeurIPS 2025 · 160 citations
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim et al.ICLR 2026 · 154 citations
- Cost-of-Pass: An Economic Framework for Evaluating Language ModelsMehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yüksekgönül et al.ICLR 2026 · 40 citations
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan et al.ICML 2026 · 31 citations
Builds on8
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 338 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn et al.NeurIPS 2023 · 81 citations
Related papers
- SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language ModelsJingxuan Xu, Ken Deng, Weihao Li, Songwei Yu et al.ICML 2026 · 9 citations
- Unified Software Engineering Agent as AI Software EngineerLeonhard Applis, Yuntong Zhang, Shanchao Liang, Nan Jiang et al.ICSE 2026
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?John Yang, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret et al.ICLR 2025
- RepoGraph: Enhancing AI Software Engineering with Repository-level Code GraphSiru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao et al.ICLR 2025
