Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning Approach
Junzhe Zhang
摘要
A dynamic treatment regime (DTR) consists of a sequence of decision rules, one per stage of intervention, that dictates how to determine the treatment assignment to patients based on evolving treatments and covariates' history. These regimes are particularly effective for managing chronic disorders and is arguably one of the critical ingredients underlying more personalized decisionmaking systems. All reinforcement learning algorithms for finding the optimal DTR in online settings will suffer Ω( |D X∪S |T ) regret on some environments, where T is the number of experiments and D X∪S is the domains of the treatments X and covariates S. This implies that T = Ω(|D X∪S |) trials will be required to generate an optimal DTR. In many applications, the domains of X and S could be enormous, which means that the time required to ensure appropriate learning may be unattainable. We show that, if the causal diagram of the underlying environment is provided, one could achieve regret that is exponentially smaller than D X∪S . In particular, we develop two online algorithms that satisfy such regret bounds by exploiting the causal structure underlying the DTR; one is the based on the principle of optimism in the face of uncertainty (OFU-DTR), and the other uses the posterior sampling learning (PS-DTR). Finally, we introduce efficient methods to accelerate these online learning procedures by leveraging the abundant, yet biased observational (non-experimental) data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 被引用 18 次
- Training a Resilient Q-network against Observational InterferenceChao-Han Huck Yang, I-Te Danny Hung, Yi Ouyang, Pin-Yu ChenAAAI 2022 · 被引用 18 次
- Online Reinforcement Learning for Mixed Policy ScopesJunzhe Zhang, Elias BareinboimNeurIPS 2022 · 被引用 11 次
- Constrained Causal Bayesian OptimizationVirginia Aglietti, Alan Malek, Ira Ktena, Silvia ChiappaICML 2023 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Gradient Regularized V-Learning for Dynamic Treatment RegimesYao Zhang, Mihaela van der SchaarNeurIPS 2020 · 被引用 6 次
- Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement LearningMirco Mutti, Riccardo De Santi, Marcello Restelli, Alexander Marx 等ICLR 2024 · 被引用 6 次
- Provably Efficient Causal Reinforcement Learning with Confounded Observational DataLingxiao Wang, Zhuoran Yang, Zhaoran WangNeurIPS 2021 · 被引用 61 次
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 被引用 7 次
- Bayesian Optimistic Optimization: Optimistic Exploration for Model-based Reinforcement LearningChenyang Wu, Tianci Li, Zongzhang Zhang, Yang YuNeurIPS 2022 · 被引用 9 次
