BRIDGE: Predicting Human Task Completion Time From Model Performance
Fengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau, Siva Reddy, Hugo Larochelle
摘要
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on direct human task completion time annotations are costly, noisy, and difficult to scale across benchmarks. In this work, we propose BRIDGE, a unified psychometric framework that learns a latent difficulty scale from model responses and anchors it to human task completion time. Using a two-parameter logistic Item Response Theory model, we jointly estimate latent task difficulty and model capability from model performance data across multiple benchmarks. We demonstrate that latent task difficulty varies linearly with the logarithm of human completion time, allowing human task completion time to be inferred for new benchmarks from model performance alone. Leveraging this alignment, we forecast frontier model capabilities in terms of human task length and independently reproduce METR's exponential scaling results, with the 50% solvable task horizon doubling approximately every 6 months. Code Repository: McGill-NLP/BRIDGE 1 Introduction As Artificial Intelligence (AI) systems are increasingly deployed in open-ended, real-world settings, their capabilities are typically reported via benchmark scores and aggregate metrics. However, such scores are difficult to interpret as measures of task difficulty or real-world effort: improvements may reflect gains on short, routine instances while leaving longer, multi-step tasks largely unchanged, and comparable score changes can correspond to very different shifts in practical capability. Consequently, benchmark performance alone provides limited guidance about what AI systems can reliably do or how quickly practical capability is improving. A more actionable evaluation should express performance on human-aligned scales, most directly, the time a human would require to complete a task. A prominent attempt to express AI capability in human terms is by Kwa et al. (2026) at METR, which measures performance in terms of the length of tasks AI agents can complete and reports exponential growth, with task-length horizons doubling roughly every 7 months. While compelling, this paradigm depends on human task completion time annotations from people with relevant expertise. These annotations are expensive collect and are difficult to extend consistently across diverse benchmarks. As model capabilities continue to scale, relying on new human studies to anchor each benchmark becomes increasingly impractical, creating a growing gap between benchmark-centric evaluation and human-centric notions of difficulty. In this work, we introduce BRIDGE, 1 a unified psychometric framework (illustrated in Figure 1 ) that addresses this gap by aligning task difficulty for humans, measured by task completion time, with task difficulty for models, measured by benchmark performance. Item Response Theory (IRT) (Baker, 2001) , and in particular the two-parameter logistic (2PL) model, has been adopted to analyze large language model (LLM) performance and benchmarks
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng 等NeurIPS 2025 · 被引用 160 次
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim 等ICLR 2026 · 被引用 154 次
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung 等ICLR 2025 · 被引用 9 次
- SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty PredictionAlexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge 等EMNLP 2025
相关 Paper
- Bridging Human and LLM Judgments: Understanding and Narrowing the GapFelipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu 等NeurIPS 2025 · 被引用 7 次
- Reliable and Efficient Amortized Model-based EvaluationSang T. Truong, Yuheng Tu, Percy Liang, Bo Li 等ICML 2025
- Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling EstimationSang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi KoyejoICML 2026
- RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsHjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar 等ICML 2025
- Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use AgentsQihao Wang, Yue Hu, Mingzhe Lu, Jiayue Wu 等AAAI 2026
