Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment Regimes
Changchang Yin, Ruoqi Liu, Jeffrey M. Caterino, Ping Zhang
摘要
Despite intense efforts in basic and clinical research, an individualized ventilation strategy for critically ill patients remains a major challenge. Recently, dynamic treatment regime (DTR) with reinforcement learning (RL) on electronic health records (EHR) has attracted interest from both the healthcare industry and machine learning research community. However, most learned DTR policies might be biased due to the existence of confounders. Although some treatment actions non-survivors received may be helpful, if confounders cause the mortality, the training of RL models guided by long-term outcomes (e.g., 90-day mortality) would punish those treatment actions causing the learned DTR policies to be suboptimal. In this study, we develop a new deconfounding actor-critic network (DAC) to learn optimal DTR policies for patients. To alleviate confounding issues, we incorporate a patient resampling module and a confounding balance module into our actor-critic framework. To avoid punishing the effective treatment actions non-survivors received, we design a short-term reward to capture patients' immediate health state changes. Combining short-term with long-term rewards could further improve the model performance. Moreover, we introduce a policy adaptation method to successfully transfer the learned model to new-source small-scale datasets. The experimental results on one semi-synthetic and two different real-world datasets show the proposed model outperforms the state-of-the-art models. The proposed model provides individualized treatment decisions for mechanical ventilation that could improve patient outcomes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisShao Zhang, Jianing Yu, Xuhai Xu, Changchang Yin 等CHI 2024 · 被引用 95 次
- medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision SupportQianyi Xu, Gousia Habib, Feng Wu, Dilruk Perera 等KDD 2026 · 被引用 4 次
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan 等AAAI 2026
它引用的顶会 Paper5
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 被引用 224 次
- Time Series Deconfounder: Estimating Treatment Effects over Time in the Presence of Hidden ConfoundersIoana Bica, Ahmed M. Alaa, Mihaela van der SchaarICML 2020 · 被引用 133 次
- Identifying Sepsis Subphenotypes via Time-Aware Multi-Modal Auto-EncoderChangchang Yin, Ruoqi Liu, Dongdong Zhang, Ping ZhangKDD 2020 · 被引用 46 次
- Training a Resilient Q-network against Observational InterferenceChao-Han Huck Yang, I-Te Danny Hung, Yi Ouyang, Pin-Yu ChenAAAI 2022 · 被引用 18 次
- Gradient Regularized V-Learning for Dynamic Treatment RegimesYao Zhang, Mihaela van der SchaarNeurIPS 2020 · 被引用 6 次
相关 Paper
- Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱Lu Wang, Wenchao Yu, Xiaofeng He, Wei Cheng 等WWW 2020 · 被引用 33 次
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 被引用 78 次
- Benchmarking Reinforcement Learning Algorithms for ICU Ventilator Settings: An Interpretable and Probabilistic Patient Environment for Doctor AgentsYa-Hsi Chang, Po-Chih KuoAAAI 2026
- Delphic Offline Reinforcement Learning under Nonidentifiable Hidden ConfoundingAlizée Pace, Hugo Yèche, Bernhard Schölkopf, Gunnar Rätsch 等ICLR 2024 · 被引用 9 次
- Confounding Robust Deep Reinforcement Learning: A Causal ApproachMingxuan Li, Junzhe Zhang, Elias BareinboimNeurIPS 2025 · 被引用 7 次
