Towards Safe Policy Learning under Partial Identifiability: A Causal Approach
Shalmali Joshi, Junzhe Zhang, Elias Bareinboim
摘要
Learning personalized treatment policies is a formative challenge in many real-world applications, including in healthcare, econometrics, artificial intelligence. However, the effectiveness of candidate policies is not always identifiable, i.e., it is not uniquely computable from the combination of the available data and assumptions about the generating mechanisms. This paper studies policy learning from data collected in various non-identifiable settings, i.e., (1) observational studies with unobserved confounding; (2) randomized experiments with partial observability; and (3) their combinations. We derive sharp, closed-formed bounds from observational and experimental data over the conditional treatment effects. Based on these novel bounds, we further characterize the problem of safe policy learning and develop an algorithm that trains a policy from data guaranteed to achieve, at least, the performance of the baseline policy currently deployed. Finally, we validate our proposed algorithm on synthetic data and a large clinical trial, demonstrating that it guarantees safe behaviors and robust performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 被引用 12 次
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 被引用 7 次
- Towards Estimating Bounds on the Effect of Policies under Unobserved ConfoundingAlexis Bellot, Silvia ChiappaNeurIPS 2024 · 被引用 6 次
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- Data Fusion for Partial Identification of Causal EffectsQuinn Lanners, Cynthia Rudin, Alexander Volfovsky, Harsh ParikhNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper5
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 被引用 81 次
- Partial Counterfactual Identification from Observational and Experimental DataJunzhe Zhang, Jin Tian, Elias BareinboimICML 2022 · 被引用 77 次
- Bounding Causal Effects on Continuous OutcomeJunzhe Zhang, Elias BareinboimAAAI 2021 · 被引用 49 次
- Causal Inference Through the Structural Causal Marginal ProblemLuigi Gresele, Julius von Kügelgen, Jonas M. Kübler, Elke Kirschbaum 等ICML 2022 · 被引用 28 次
- Causal Effect Identifiability under Partial-ObservabilitySanghack Lee, Elias BareinboimICML 2020 · 被引用 26 次
相关 Paper
- Confounding-Robust Policy Evaluation in Infinite-Horizon Reinforcement LearningNathan Kallus, Angela ZhouNeurIPS 2020 · 被引用 78 次
- Meta-Learners for Partially-Identified Treatment Effects Across Multiple EnvironmentsJonas Schweisthal, Dennis Frauen, Mihaela van der Schaar, Stefan FeuerriegelICML 2024 · 被引用 10 次
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 被引用 7 次
- Learning Treatment Allocations with Risk Control Under Partial IdentifiabilitySofia Ek, Dave ZachariahICML 2026
- Estimation of Bounds on Potential Outcomes For Decision MakingMaggie Makar, Fredrik D. Johansson, John V. Guttag, David A. SontagICML 2020 · 被引用 11 次
