Adversarial Cooperative Imitation Learning for Dynamic Treatment Regimes✱
Lu Wang, Wenchao Yu, Xiaofeng He, Wei Cheng, Martin Renqiang Ren, Wei Wang, Bo Zong, Haifeng Chen, Hongyuan Zha
摘要
Recent developments in discovering dynamic treatment regimes (DTRs) have heightened the importance of deep reinforcement learning (DRL) which are used to recover the doctor’s treatment policies. However, existing DRL-based methods expose the following limitations: 1) supervised methods based on behavior cloning suffer from compounding errors; 2) the self-defined reward signals in reinforcement learning models are either too sparse or need clinical guidance; 3) only positive trajectories (e.g. survived patients) are considered in current imitation learning models, with negative trajectories (e.g. deceased patients) been largely ignored, which are examples of what not to do and could help the learned policy avoid repeating mistakes. To address these limitations, in this paper, we propose the adversarial cooperative imitation learning model, ACIL, to deduce the optimal dynamic treatment regimes that mimics the positive trajectories while differs from the negative trajectories. Specifically, two discriminators are used to help achieve this goal: an adversarial discriminator is designed to minimize the discrepancies between the trajectories generated from the policy and the positive trajectories, and a cooperative discriminator is used to distinguish the negative trajectories from the positive and generated trajectories. The reward signals from the discriminators are utilized to refine the policy for dynamic treatment regimes. Experiments on the publicly real-world medical data demonstrate that ACIL improves the likelihood of patient survival and provides better dynamic treatment regimes with the exploitation of information from both positive and negative trajectories.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Causal Imitation for Markov Decision Processes: a Partial Identification ApproachKangrui Ruan, Junzhe Zhang, Xuan Di, Elias BareinboimNeurIPS 2024 · 被引用 12 次
- Inverse Factorized Soft Q-Learning for Cooperative Multi-agent Imitation LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 被引用 10 次
- FedSkill: Privacy Preserved Interpretable Skill Learning via ImitationYushan Jiang, Wenchao Yu, Dongjin Song, Lu Wang 等KDD 2023 · 被引用 5 次
- Skill Disentanglement for Imitation Learning from Suboptimal DemonstrationsTianxiang Zhao, Wenchao Yu, Suhang Wang, Lu Wang 等KDD 2023 · 被引用 5 次
相关 Paper
- Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment RegimesChangchang Yin, Ruoqi Liu, Jeffrey M. Caterino, Ping ZhangKDD 2022 · 被引用 5 次
- Policy Contrastive Imitation LearningJialei Huang, Zhao-Heng Yin, Yingdong Hu, Yang GaoICML 2023 · 被引用 4 次
- Unlabeled Imperfect Demonstrations in Adversarial Imitation LearningYunke Wang, Bo Du, Chang XuAAAI 2023 · 被引用 11 次
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous DrivingXiaoji Zheng, Ziyuan Yang, Yanhao Chen, Yuhang PENG 等ICML 2026
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
