Prediction-Powered Causal Inferences
Riccardo Cadei, Ilker Demirel, Piersilvio De Bartolomeis, Lukas Lindorfer, Sylvia Cremer, Cordelia Schmid, Francesco Locatello
Abstract
In many scientific experiments, the data annotating cost constraints the pace for testing novel hypotheses. Yet, modern machine learning pipelines offer a promising solution, provided their predictions yield correct conclusions. We focus on Prediction-Powered Causal Inferences (PPCI), i.e., estimating the treatment effect in an unlabeled target experiment, relying on training data with the same outcome annotated but potentially different treatment or effect modifiers. We first show that conditional calibration guarantees valid PPCI at population level. Then, we introduce a sufficient representation constraint transferring validity across experiments, which we propose to enforce in practice in Deconfounded Empirical Risk Minimization, our new model-agnostic training objective. We validate our method on synthetic and real-world scientific data, solving impossible problem instances for Empirical Risk Minimization even with standard invariance constraints. In particular, for the first time, we achieve valid causal inference on a scientific experiment with complex recording and no human annotations, fine-tuning a foundational model on our similar annotated experiment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66897bf3-a5d5-4579-a688-5a4e83bed8e9Cited by top-tier papers2
- The third pillar of causal analysis? A measurement perspective on causal representationsDingling Yao, Shimeng Huang, Riccardo Cadei, Kun Zhang et al.NeurIPS 2025 · 5 citations
- Exploratory Causal Inference in SAEnceTommaso Mencattini, Riccardo Cadei, Francesco LocatelloICLR 2026 · 4 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
Related papers
- Partial Transportability for Domain GeneralizationKasra Jalaldoust, Alexis Bellot, Elias BareinboimNeurIPS 2024 · 14 citations
- MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean EstimationSe Yoon Lee, Jae-kwang KimICML 2026 · 1 citation
- Prediction-Powered Adaptive Shrinkage EstimationSida Li, Nikolaos IgnatiadisICML 2025
- Generative multitask learning mitigates target-causing confoundingTaro Makino, Krzysztof J. Geras, Kyunghyun ChoNeurIPS 2022 · 9 citations
- Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc RegressionBenjamin Eyre, David MadrasICML 2025
