ICML2026
CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?
Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha, Xianlin Sun, Pepijn Cobben, Maximilian Mordig, Jacob Emmerson, Anahita Haghighat, Furkan Danisman, Yuen Chen, Clijo Jose, Andrei Muresanu, Justin Cui, Jiarui Liu, Yahang Qi, Punya Pandey, Yinya Huang, Bernhard Schölkopf, Zhijing Jin
Abstract
Identifying and estimating causal relationships from data is a crucial component of empirical research. While large language model-powered tools have shown potential for assisting research workflows, their ability to perform end-to-end causal inference remains underexplored. We introduce CauSciBench, a benchmark that puts LLM-powered tools to the test on causality- driven research questions. Unlike previous related benchmarks that focus on coding alone, CauSciBench enables evaluation across the full pipeline of causal inference: from method and variable selection to computation of causal effects and statistical interpretation in the context of real-world research problems. We evaluated 7 frontier models on over 300 queries derived from scientific publications, textbook problems, sem- inal datasets, and synthetic scenarios. Results show that models consistently perform worse on real datasets, with the key bottleneck being the selection of an appropriate causal inference method.