GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simón Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Samuel Miserendino
摘要
We introduce GDPval, a benchmark evaluating AI model capabilities on realworld economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional ReasoningAfra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta 等ACL 2026 · 被引用 18 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
- BRIDGE: Predicting Human Task Completion Time From Model PerformanceFengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau 等ICML 2026 · 被引用 5 次
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu 等CHI 2026 · 被引用 1 次
- A Framework to Characterize Reporting on Generative AI UseAgathe Balayn, Varun Nagaraj Rao, Su Lin Blodgett, Aylin Caliskan 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper4
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?Samuel Miserendino, Michele Wang, Tejal Patwardhan, Johannes HeideckeICML 2025
相关 Paper
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng 等NeurIPS 2025 · 被引用 160 次
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao 等ICML 2026
- RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsHjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar 等ICML 2025
- Expanding the AI Evaluation Toolbox with Statistical ModelsDrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao 等ICML 2026 · 被引用 4 次
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and JudgeZhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao 等ICLR 2026 · 被引用 20 次
