Lune

NeurIPS2021顶会

The Causal-Neural Connection: Expressiveness, Learnability, and Inference

Kevin Xia, Kai-Zhan Lee, Yoshua Bengio, Elias Bareinboim

2021年份
158被引次数
41顶会引用

摘要

One of the central elements of any causal inference is an object called structural causal model (SCM), which represents a collection of mechanisms and exogenous sources of random variation of the system under investigation (Pearl, 2000) . An important property of many kinds of neural networks is universal approximability: the ability to approximate any function to arbitrary precision. Given this property, one may be tempted to surmise that a collection of neural nets is capable of learning any SCM by training on data generated by that SCM. In this paper, we show this is not the case by disentangling the notions of expressivity and learnability. Specifically, we show that the causal hierarchy theorem (Thm. 1, Bareinboim et al., 2020) , which describes the limits of what can be learned from data, still holds for neural models. For instance, an arbitrarily complex and expressive neural net is unable to predict the effects of interventions given observational data alone. Given this result, we introduce a special type of SCM called a neural causal model (NCM), and formalize a new type of inductive bias to encode structural constraints necessary for performing causal inferences. Building on this new class of models, we focus on solving two canonical tasks found in the literature known as causal identification and estimation. Leveraging the neural toolbox, we develop an algorithm that is both sufficient and necessary to determine whether a causal effect can be learned from data (i.e., causal identifiability); it then estimates the effect whenever identifiability holds (causal estimation). Simulations corroborate the proposed approach. 1 This structure is named after Judea Pearl and is a central topic in his Book of Why (BoW), where it is also called the "Ladder of Causation" [59] . For a more technical discussion on the PCH, we refer readers to [5] . 2 The full inferential challenge is, in practice, more general since an agent may be able to perform interventions and obtain samples from a subset of the PCH's layers, while its goal is to make inferences about some other parts of the layers [7, 46, 5] . This situation is not uncommon in RL settings [66, 17, 44, 45] . Still, for the sake of space and concreteness, we will focus on two canonical and more basic tasks found in the literature. 3 We defer a more formal discussion on how neural models could be used to assess the effect of interventions to Sec. 2. Still, this is neither attainable in all universal neural architectures nor trivially implementable. 4 Pearl shared a similar observation in the BoW [59, p. 32]: "Without the causal model, we could not go from rung (layer) one to rung (layer) two. This is why deep-learning systems (as long as they use only rung-one data and do not have a causal model) will never be able to answer questions about interventions (...)".

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper41

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖