CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning
Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas CHESNEAU, Marc Schoenauer, Özgür Şimşek
Abstract
Causal machine learning aims to answer "what if" questions using machine learning algorithms, making it a promising tool for high-stakes decision-making. Yet, empirical evaluation practices remain limited. Existing benchmarks often rely on a handful of hand-crafted or semi-synthetic datasets, leading to brittle, non-generalizable conclusions. To bridge this gap, we introduce CausalProfiler, a synthetic benchmark generator for causal machine learning methods. Based on a set of explicit design choices about the class of causal models, queries, and data considered, CausalProfiler randomly samples causal models, data, queries, and ground truths constituting the synthetic causal benchmarks. In this way, causal machine learning can be rigorously and transparently evaluated under a variety of conditions. This work offers the first random generator of synthetic causal benchmarks with coverage guarantees and transparent assumptions operating on the three levels of causal reasoning: observation, intervention, and counterfactual. We demonstrate its utility by evaluating several state-of-the-art methods under diverse conditions and assumptions, both in and out of the identification regime, illustrating the types of analyses and insights CausalProfiler enables.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
Related papers
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha et al.ICML 2026
- The robustness of differentiable Causal Discovery in misspecified ScenariosHuiyang Yi, Yanyan He, Duxin Chen, Mingyu Kang et al.ICLR 2025
- Deep Counterfactual Estimation with Categorical Background VariablesEdward De BrouwerNeurIPS 2022 · 8 citations
- Language Models as Causal Effect GeneratorsLucius E. J. Bynum, Kyunghyun ChoEMNLP 2025
