Lune

ICML2025Top-tier venue

SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes Heidecke

2025Year
18Top-tier citations

Abstract

We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at 1millionUSDtotalinrealworldpayouts.SWE−Lancerencompassesbothindependentengineeringtasks−rangingfrom1 million USD total in realworld payouts. SWE-Lancer encompasses both independent engineering tasks -ranging from 50 bug fixes to $32,000 feature implementations -and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with endto-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond ( https://github.com/ openai/SWELancer-Benchmark ). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c78c709c-c791-4ae4-88f6-733d9bff2f8b

Cited by top-tier papers18

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines