Smoke and Mirrors in Causal Downstream Tasks
Riccardo Cadei, Lukas Lindorfer, Sylvia Cremer, Cordelia Schmid, Francesco Locatello
摘要
Machine Learning and AI have the potential to transform data-driven scientific discovery, enabling accurate predictions for several scientific phenomena. As many scientific questions are inherently causal, this paper looks at the causal inference task of treatment effect estimation, where the outcome of interest is recorded in high-dimensional observations in a Randomized Controlled Trial (RCT). Despite being the simplest possible causal setting and a perfect fit for deep learning, we theoretically find that many common choices in the literature may lead to biased estimates. To test the practical impact of these considerations, we recorded ISTAnt, the first real-world benchmark for causal inference downstream tasks on high-dimensional observations as an RCT studying how garden ants (Lasius neglectus) respond to microparticles applied onto their colony members by hygienic grooming. Comparing 6 480 models fine-tuned from state-of-the-art visual backbones, we find that the sampling and modeling choices significantly affect the accuracy of the causal estimate, and that classification accuracy is not a proxy thereof. We further validated the analysis, repeating it on a synthetically generated visual data set controlling the causal model. Our results suggest that future benchmarks should carefully consider real downstream scientific questions, especially causal ones. Further, we highlight guidelines for representation learning methods to help answer causal questions in the sciences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Prediction-Powered Causal InferencesRiccardo Cadei, Ilker Demirel, Piersilvio De Bartolomeis, Lukas Lindorfer 等NeurIPS 2025 · 被引用 9 次
- The third pillar of causal analysis? A measurement perspective on causal representationsDingling Yao, Shimeng Huang, Riccardo Cadei, Kun Zhang 等NeurIPS 2025 · 被引用 5 次
- Exploratory Causal Inference in SAEnceTommaso Mencattini, Riccardo Cadei, Francesco LocatelloICLR 2026 · 被引用 4 次
- Unifying Causal Representation Learning with the Invariance PrincipleDingling Yao, Dario Rancati, Riccardo Cadei, Marco Fumero 等ICLR 2025
- CaTs and DAGs: Integrating Directed Acyclic Graphs with Transformers for Causally Constrained PredictionsMatthew J. Vowels, Mathieu Rochat, Sina AkbariICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf 等ICML 2020 · 被引用 361 次
- Weakly supervised causal representation learningJohann Brehmer, Pim de Haan, Phillip Lippe, Taco S. CohenNeurIPS 2022 · 被引用 196 次
相关 Paper
- CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?Sawal Acharya, Terry J Zhang, Andrew Kim, Rahul B Shrestha 等ICML 2026
- Task-specific experimental design for treatment effect estimationBethany Connolly, Kim Moore, Tobias Schwedes, Alexander Adam 等ICML 2023 · 被引用 4 次
- Causal Representation Learning for Instantaneous and Temporal Effects in Interactive SystemsPhillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano 等ICLR 2023
- Causal Relational LearningBabak Salimi, Harsh Parikh, Moe Kayali, Lise Getoor 等SIGMOD 2020 · 被引用 38 次
- Estimating Average Causal Effects from Patient TrajectoriesDennis Frauen, Tobias Hatt, Valentyn Melnychuk, Stefan FeuerriegelAAAI 2023 · 被引用 34 次
