Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning Approach
Junzhe Zhang
Abstract
A dynamic treatment regime (DTR) consists of a sequence of decision rules, one per stage of intervention, that dictates how to determine the treatment assignment to patients based on evolving treatments and covariates' history. These regimes are particularly effective for managing chronic disorders and is arguably one of the critical ingredients underlying more personalized decisionmaking systems. All reinforcement learning algorithms for finding the optimal DTR in online settings will suffer Ω( |D X∪S |T ) regret on some environments, where T is the number of experiments and D X∪S is the domains of the treatments X and covariates S. This implies that T = Ω(|D X∪S |) trials will be required to generate an optimal DTR. In many applications, the domains of X and S could be enormous, which means that the time required to ensure appropriate learning may be unattainable. We show that, if the causal diagram of the underlying environment is provided, one could achieve regret that is exponentially smaller than D X∪S . In particular, we develop two online algorithms that satisfy such regret bounds by exploiting the causal structure underlying the DTR; one is the based on the principle of optimism in the face of uncertainty (OFU-DTR), and the other uses the posterior sampling learning (PS-DTR). Finally, we introduce efficient methods to accelerate these online learning procedures by leveraging the abundant, yet biased observational (non-experimental) data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdf5ac6b-413f-4fa1-9d71-4773e8871713Cited by top-tier papers18
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 18 citations
- Training a Resilient Q-network against Observational InterferenceChao-Han Huck Yang, I-Te Danny Hung, Yi Ouyang, Pin-Yu ChenAAAI 2022 · 18 citations
- Online Reinforcement Learning for Mixed Policy ScopesJunzhe Zhang, Elias BareinboimNeurIPS 2022 · 11 citations
- Constrained Causal Bayesian OptimizationVirginia Aglietti, Alan Malek, Ira Ktena, Silvia ChiappaICML 2023 · 9 citations
Builds on1
Related papers
- Gradient Regularized V-Learning for Dynamic Treatment RegimesYao Zhang, Mihaela van der SchaarNeurIPS 2020 · 6 citations
- Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement LearningMirco Mutti, Riccardo De Santi, Marcello Restelli, Alexander Marx et al.ICLR 2024 · 6 citations
- Provably Efficient Causal Reinforcement Learning with Confounded Observational DataLingxiao Wang, Zhuoran Yang, Zhaoran WangNeurIPS 2021 · 61 citations
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 7 citations
- Bayesian Optimistic Optimization: Optimistic Exploration for Model-based Reinforcement LearningChenyang Wu, Tianci Li, Zongzhang Zhang, Yang YuNeurIPS 2022 · 9 citations
