Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued Policies
Shengpu Tang, Aditya Modi, Michael W. Sjoding, Jenna Wiens
摘要
Standard reinforcement learning (RL) aims to find an optimal policy that identifies the best action for each state. However, in healthcare settings, many actions may be near-equivalent with respect to the reward (e.g., survival). We consider an alternative objective -- learning set-valued policies to capture near-equivalent actions that lead to similar cumulative rewards. We propose a model-free algorithm based on temporal difference learning and a near-greedy heuristic for action selection. We analyze the theoretical properties of the proposed algorithm, providing optimality guarantees and demonstrate our approach on simulated environments and a real clinical task. Empirically, the proposed algorithm exhibits good convergence properties and discovers meaningful near-equivalent actions. Our work provides theoretical, as well as practical, foundations for clinician/human-in-the-loop decision making, in which humans (e.g., clinicians, patients) can incorporate additional knowledge (e.g., side effects, patient preference) when selecting among near-equivalent actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez 等NeurIPS 2022 · 被引用 63 次
- Medical Dead-ends and Learning to Identify High-Risk States and TreatmentsMehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh GhassemiNeurIPS 2021 · 被引用 51 次
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPsHarsh Satija, Philip S. Thomas, Joelle Pineau, Romain LarocheNeurIPS 2021 · 被引用 30 次
- Counterfactual-Augmented Importance Sampling for Semi-Offline Policy EvaluationShengpu Tang, Jenna WiensNeurIPS 2023 · 被引用 8 次
- User-Interactive Offline Reinforcement LearningPhillip Swazinna, Steffen Udluft, Thomas A. RunklerICLR 2023 · 被引用 3 次
相关 Paper
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 被引用 7 次
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 被引用 18 次
- Quasi-optimal Reinforcement Learning with Continuous ActionsYuhan Li, Wenzhuo Zhou, Ruoqing ZhuICLR 2023 · 被引用 2 次
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 被引用 6 次
- Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao Luo, Ke LiICML 2026 · 被引用 3 次
