GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simón Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Samuel Miserendino
Abstract
We introduce GDPval, a benchmark evaluating AI model capabilities on realworld economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4c9ef9b-f7ca-4a7a-b808-9952869d04a5Cited by top-tier papers8
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional ReasoningAfra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta et al.ACL 2026 · 18 citations
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
- BRIDGE: Predicting Human Task Completion Time From Model PerformanceFengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau et al.ICML 2026 · 5 citations
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu et al.CHI 2026 · 1 citation
- A Framework to Characterize Reporting on Generative AI UseAgathe Balayn, Varun Nagaraj Rao, Su Lin Blodgett, Aylin Caliskan et al.CHI 2026 · 1 citation
Builds on4
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes HeideckeICML 2025
Related papers
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng et al.NeurIPS 2025 · 160 citations
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao et al.ICML 2026
- RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsHjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar et al.ICML 2025
- Expanding the AI Evaluation Toolbox with Statistical ModelsDrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao et al.ICML 2026 · 4 citations
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and JudgeZhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao et al.ICLR 2026 · 20 citations
