Gradient Regularized V-Learning for Dynamic Treatment Regimes
Yao Zhang, Mihaela van der Schaar
Abstract
Deciding how to optimally treat a patient, including how to select treatments over time among the multiple available treatments, represents one of the most important issues that need to be addressed in medicine today. A dynamic treatment regime (DTR) is a sequence of treatment rules indicating how to individualize treatments for a patient based on the previously assigned treatments and the evolving covariate history. However, DTR evaluation and learning based on offline data remain challenging problems due to the bias introduced by time-varying confounders that affect treatment assignment over time; this may lead to suboptimal treatment rules being used in practice. In this paper, we introduce Gradient Regularized V -learning (GRV), a novel method for estimating the value function of a DTR. GRV regularizes the underlying outcome and propensity score models with respect to the optimality condition in semiparametric estimation theory. On the basis of this design, we construct estimators that are efficient and stable in finite samples regime. Using multiple simulation studies and one real-world medical dataset, we demonstrate that our method is superior in DTR evaluation and learning, thereby providing improved treatment options over time for patients.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 43 citations
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 26 citations
- Efficient Policy Learning from Surrogate-Loss Classification ReductionsAndrew Bennett, Nathan KallusICML 2020 · 20 citations
Related papers
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 78 citations
- Evaluating and Learning Optimal Dynamic Treatment Regimes under Truncation by DeathSihyung Park, Wenbin Lu, Shu YangNeurIPS 2025 · 1 citation
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 7 citations
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 224 citations
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 22 citations
