Lune

ICML2026顶会

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric Xing

2026年份
2被引次数

摘要

Autonomous AI agents powered by large language models (LLMs) are increasingly capable of running a full cycle of scientific research, yet we still lack reliable ways to verify that their discoveries are correct. Because novel findings demand costly real-world validation, existing benchmarks fall back on LLM-as-judge scoring of generated papers or single leaderboard metrics, both coarse proxies for scientific reasoning. We introduce FIRE-BENCH (Full-cycle Insight Rediscovery Evaluation), which instead asks agents to rediscover established, verifiable findings from recent, high-impact machine learning research. Given only a high-level research question from a published study, an agent must independently design experiments, run them, and draw evidence-backed conclusions, scored against the study's documented findings. Across state-of-the-art agents with frontier backbones such as gpt-5, even the strongest reaches limited rediscovery success (<50 F1), with high run-to-run variance and recurring failures in experimental design, execution, and evidence-based reasoning. Beyond diagnosing current systems, FIRE-BENCH shows that open-ended discovery can be evaluated rigorously and verifiably, laying a foundation for building reliable environments that improve agents.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper39

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖