SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
Udari Sehwag, Elaine Lau, Haniyeh Oskouie, Shayan Shabihi, Erich Liang, Andrea Toledo, Guillermo Mangialardi, Sergio Fonrouge, Ed-Yeremai Hernandez-Cardona, Paula Vergara, Utkarsh Tyagi, Chen Bo Calvin Zhang
Abstract
Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their ability to predict experimental outcomes---a task where AI could significantly exceed human capabilities---remains largely underexplored. We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry. SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process? Evaluations reveal fundamental limitations on both fronts. Model accuracies are 14-26% and human expert performance is 20%. Although some frontier models exceed human performance model accuracy is still far below what would enable reliable experimental guidance. Even within the limited performance, models fail to distinguish reliable predictions from unreliable ones, achieving only 20% accuracy regardless of their confidence or whether they judge outcomes as predictable without physical experimentation. Human experts, in contrast, demonstrate strong calibration: their accuracy increases from 5% to 80% as they deem outcomes more predictable without conducting the experiment. SciPredict establishes a rigorous framework demonstrating that superhuman performance in experimental science requires not just better predictions, but better awareness of prediction reliability. For reproducibility all our data and code are provided at https://anonymous.4open.science/r/SciPredict-AI01.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 620f0c36-7e2e-4433-ad81-e1bda800b7a5Builds on9
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
- SR-Scientist: Scientific Equation Discovery With Agentic AIShijie Xia, Yuhan Sun, Pengfei LiuICLR 2026 · 28 citations
- AutoReproduce: Automatic AI Experiment Reproduction with Paper LineageXuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi et al.ACL 2026 · 18 citations
Related papers
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani et al.ICML 2025
- SCI-Verifier: Scientific Verifier with ThinkingShenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao et al.ICLR 2026 · 5 citations
- CURIE: Evaluating LLMs on Multitask Scientific Long-Context Understanding and ReasoningHao Cui, Zahra Shamsi, Gowoon Cheon, Xuejian Ma et al.ICLR 2025
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsXin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen et al.ICML 2025
