Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment Regimes
Changchang Yin, Ruoqi Liu, Jeffrey M. Caterino, Ping Zhang
Abstract
Despite intense efforts in basic and clinical research, an individualized ventilation strategy for critically ill patients remains a major challenge. Recently, dynamic treatment regime (DTR) with reinforcement learning (RL) on electronic health records (EHR) has attracted interest from both the healthcare industry and machine learning research community. However, most learned DTR policies might be biased due to the existence of confounders. Although some treatment actions non-survivors received may be helpful, if confounders cause the mortality, the training of RL models guided by long-term outcomes (e.g., 90-day mortality) would punish those treatment actions causing the learned DTR policies to be suboptimal. In this study, we develop a new deconfounding actor-critic network (DAC) to learn optimal DTR policies for patients. To alleviate confounding issues, we incorporate a patient resampling module and a confounding balance module into our actor-critic framework. To avoid punishing the effective treatment actions non-survivors received, we design a short-term reward to capture patients' immediate health state changes. Combining short-term with long-term rewards could further improve the model performance. Moreover, we introduce a policy adaptation method to successfully transfer the learned model to new-source small-scale datasets. The experimental results on one semi-synthetic and two different real-world datasets show the proposed model outperforms the state-of-the-art models. The proposed model provides individualized treatment decisions for mechanical ventilation that could improve patient outcomes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5407294c-e324-4a34-9a5a-ed449760ce96Cited by top-tier papers3
- Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisShao Zhang, Jianing Yu, Xuhai Xu, Changchang Yin et al.CHI 2024 · 95 citations
- medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision SupportQianyi Xu, Gousia Habib, Feng Wu, Dilruk Perera et al.KDD 2026 · 4 citations
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan et al.AAAI 2026
Builds on5
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 224 citations
- Time Series Deconfounder: Estimating Treatment Effects over Time in the Presence of Hidden ConfoundersIoana Bica, Ahmed M. Alaa, Mihaela van der SchaarICML 2020 · 133 citations
- Identifying Sepsis Subphenotypes via Time-Aware Multi-Modal Auto-EncoderChangchang Yin, Ruoqi Liu, Dongdong Zhang, Ping ZhangKDD 2020 · 46 citations
- Training a Resilient Q-network against Observational InterferenceChao-Han Huck Yang, I-Te Danny Hung, Yi Ouyang, Pin-Yu ChenAAAI 2022 · 18 citations
- Gradient Regularized V-Learning for Dynamic Treatment RegimesYao Zhang, Mihaela van der SchaarNeurIPS 2020 · 6 citations
Related papers
- Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱Lu Wang, Wenchao Yu, Xiaofeng He, Wei Cheng et al.WWW 2020 · 33 citations
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 78 citations
- Benchmarking Reinforcement Learning Algorithms for ICU Ventilator Settings: An Interpretable and Probabilistic Patient Environment for Doctor AgentsYa-Hsi Chang, Po-Chih KuoAAAI 2026
- Delphic Offline Reinforcement Learning under Nonidentifiable Hidden ConfoundingAlizée Pace, Hugo Yèche, Bernhard Schölkopf, Gunnar Rätsch et al.ICLR 2024 · 9 citations
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 7 citations
