Policy-Based Bayesian Active Causal Discovery with Deep Reinforcement Learning
Heyang Gao, Zexu Sun, Hao Yang, Xu Chen
摘要
Causal discovery with observational and interventional data plays an important role in numerous fields. Due to the costly and potentially risky nature of intervention experiments, selecting informative interventions is critical in real-world situations. Several recent works introduce Bayesian active learning to select interventions that maximize the expected information gain about the underlying causal relationship at each optimization step. However, there are still some limitations within these methods: (1) Local optimality. With multiple intervention experiments, selecting optimal intervention myopically at each step may drop into the local optimal point. (2) Expensive time cost. Optimizing the most informative intervention at each step is time-consuming and not suitable for adaptive experiments with strict inference speed requirements. In this study, we propose a novel method called Reinforcement Learning-based Causal Bayesian Experimental Design (RL-CBED) to reduce the risk of local optimality and accelerate intervention selection inference. Specifically, we formulate the active causal discovery problem as a partially observable Markov decision process (POMDP). We design an information gain-based sparse reward function and then improve it to a dense reward function, providing fine-grained feedback to help the RL policy learn more quickly in complex environments. Moreover, we theoretically prove that the Q-function estimator can be learned using only trajectories sampled from the prior, which can significantly reduce the time cost of training process, enabling the real-world application of our method. Extensive experiments on both synthetic and real world-inspired semi-synthetic datasets demonstrate the effectiveness of our proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Interventions, Where and How? Experimental Design for Causal Models at ScalePanagiotis Tigas, Yashas Annadani, Andrew Jesson, Bernhard Schölkopf 等NeurIPS 2022 · 被引用 68 次
- Differentiable Multi-Target Causal Bayesian Experimental DesignPanagiotis Tigas, Yashas Annadani, Desi R. Ivanova, Andrew Jesson 等ICML 2023 · 被引用 15 次
- Sample Efficient Bayesian Learning of Causal Graphs from InterventionsZihan Zhou, Muhammad Qasim Elahi, Murat KocaogluNeurIPS 2024 · 被引用 6 次
- Active Bayesian Causal InferenceChristian Toth, Lars Lorch, Christian Knoll, Andreas Krause 等NeurIPS 2022 · 被引用 52 次
- Reinforcement Causal Structure Learning on Order GraphDezhi Yang, Guoxian Yu, Jun Wang, Zhengtian Wu 等AAAI 2023 · 被引用 20 次
