CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?
Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha, Xianlin Sun, Pepijn Cobben, Maximilian Mordig, Jacob Emmerson, Anahita Haghighat, Furkan Danisman, Yuen Chen, Clijo Jose
Abstract
Identifying and estimating causal relationships from data is a crucial component of empirical research. While large language model-powered tools have shown potential for assisting research workflows, their ability to perform end-to-end causal inference remains underexplored. We introduce CauSciBench, a benchmark that puts LLM-powered tools to the test on causality- driven research questions. Unlike previous related benchmarks that focus on coding alone, CauSciBench enables evaluation across the full pipeline of causal inference: from method and variable selection to computation of causal effects and statistical interpretation in the context of real-world research problems. We evaluated 7 frontier models on over 300 queries derived from scientific publications, textbook problems, sem- inal datasets, and synthetic scenarios. Results show that models consistently perform worse on real datasets, with the key bottleneck being the selection of an appropriate causal inference method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff et al.ICLR 2024 · 186 citations
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment TasksAllen Nie, Yuhui Zhang, Atharva Amdekar, Chris Piech et al.NeurIPS 2023 · 78 citations
Related papers
- Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal InferenceJin Du, Li Chen, Xun Xian, An Luo et al.ICLR 2026 · 4 citations
- Can Large Language Models Infer Causal Relationships from Real-World Text?Ryan Saklad, Aman Chadha, Oleg V. Pavlov, Raha MoraffahACL 2026 · 4 citations
- DiscoveryBench: Towards Data-Driven Discovery with Large Language ModelsBodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra et al.ICLR 2025
- REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic TasksLinna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen et al.AAAI 2026
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 3 citations
