The Causal-Neural Connection: Expressiveness, Learnability, and Inference
Kevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias Bareinboim
Abstract
One of the central elements of any causal inference is an object called structural causal model (SCM), which represents a collection of mechanisms and exogenous sources of random variation of the system under investigation (Pearl, 2000) . An important property of many kinds of neural networks is universal approximability: the ability to approximate any function to arbitrary precision. Given this property, one may be tempted to surmise that a collection of neural nets is capable of learning any SCM by training on data generated by that SCM. In this paper, we show this is not the case by disentangling the notions of expressivity and learnability. Specifically, we show that the causal hierarchy theorem (Thm. 1, Bareinboim et al., 2020) , which describes the limits of what can be learned from data, still holds for neural models. For instance, an arbitrarily complex and expressive neural net is unable to predict the effects of interventions given observational data alone. Given this result, we introduce a special type of SCM called a neural causal model (NCM), and formalize a new type of inductive bias to encode structural constraints necessary for performing causal inferences. Building on this new class of models, we focus on solving two canonical tasks found in the literature known as causal identification and estimation. Leveraging the neural toolbox, we develop an algorithm that is both sufficient and necessary to determine whether a causal effect can be learned from data (i.e., causal identifiability); it then estimates the effect whenever identifiability holds (causal estimation). Simulations corroborate the proposed approach. 1 This structure is named after Judea Pearl and is a central topic in his Book of Why (BoW), where it is also called the "Ladder of Causation" [59] . For a more technical discussion on the PCH, we refer readers to [5] . 2 The full inferential challenge is, in practice, more general since an agent may be able to perform interventions and obtain samples from a subset of the PCH's layers, while its goal is to make inferences about some other parts of the layers [7, 46, 5] . This situation is not uncommon in RL settings [66, 17, 44, 45] . Still, for the sake of space and concreteness, we will focus on two canonical and more basic tasks found in the literature. 3 We defer a more formal discussion on how neural models could be used to assess the effect of interventions to Sec. 2. Still, this is neither attainable in all universal neural architectures nor trivially implementable. 4 Pearl shared a similar observation in the BoW [59, p. 32]: "Without the causal model, we could not go from rung (layer) one to rung (layer) two. This is why deep-learning systems (as long as they use only rung-one data and do not have a causal model) will never be able to answer questions about interventions (...)".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1fc3209-aca2-45eb-9e24-7da3213f9124Cited by top-tier papers41
- High Fidelity Image Counterfactuals with Probabilistic Causal ModelsFabio De Sousa Ribeiro, Tian Xia, Miguel Monteiro, Nick Pawlowski et al.ICML 2023 · 68 citations
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 62 citations
- Unicorn: reasoning about configurable system performance through the lens of causalityMd Shahriar Iqbal, Rahul Krishna, Mohammad Ali Javidian, Baishakhi Ray et al.EuroSys 2022 · 60 citations
- CausalPFN: Amortized Causal Effect Estimation via In-Context LearningVahid Balazadeh Meresht, Hamidreza Kamkari, Valentin Thomas, Junwei Ma et al.NeurIPS 2025 · 52 citations
- Counterfactual Identifiability of Bijective Causal ModelsArash Nasr-Esfahany, Mohammad Alizadeh, Devavrat ShahICML 2023 · 42 citations
Builds on9
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien et al.NeurIPS 2020 · 295 citations
- Causal Discovery from Soft Interventions with Unknown Targets: Characterization and LearningAmin Jaber, Murat Kocaoglu, Karthikeyan Shanmugam, Elias BareinboimNeurIPS 2020 · 136 citations
- DeepMatch: Balancing Deep Covariate Representations for Causal Inference Using Adversarial TrainingNathan KallusICML 2020 · 84 citations
- Estimating Identifiable Causal Effects through Double Machine LearningYonghan Jung, Jin Tian, Elias BareinboimAAAI 2021 · 70 citations
Related papers
- A Generative Adversarial Framework for Bounding Confounded Causal EffectsYaowei Hu, Yongkai Wu, Lu Zhang, Xintao WuAAAI 2021 · 32 citations
- Exogenous Isomorphism for Counterfactual IdentifiabilityYikang Chen, Dehui DuICML 2025
- Relational Structural Causal ModelsAdiba Ejaz, Elias BareinboimICML 2026
- Towards Learning and Explaining Indirect Causal Effects in Neural NetworksAbbavaram Gowtham Reddy, Saketh Bachu, Harsharaj Pathak, Benin Godfrey L et al.AAAI 2024 · 3 citations
- Causal Discovery and Inference through Next-Token PredictionEivinas Butkus, Nikolaus KriegeskorteNeurIPS 2025 · 3 citations
