Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued Policies
Shengpu Tang, Aditya Modi, Michael W. Sjoding, Jenna Wiens
Abstract
Standard reinforcement learning (RL) aims to find an optimal policy that identifies the best action for each state. However, in healthcare settings, many actions may be near-equivalent with respect to the reward (e.g., survival). We consider an alternative objective -- learning set-valued policies to capture near-equivalent actions that lead to similar cumulative rewards. We propose a model-free algorithm based on temporal difference learning and a near-greedy heuristic for action selection. We analyze the theoretical properties of the proposed algorithm, providing optimality guarantees and demonstrate our approach on simulated environments and a real clinical task. Empirically, the proposed algorithm exhibits good convergence properties and discovers meaningful near-equivalent actions. Our work provides theoretical, as well as practical, foundations for clinician/human-in-the-loop decision making, in which humans (e.g., clinicians, patients) can incorporate additional knowledge (e.g., side effects, patient preference) when selecting among near-equivalent actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 553d4dcd-ffc8-476a-9e1c-0292d4362e84Cited by top-tier papers6
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez et al.NeurIPS 2022 · 63 citations
- Medical Dead-ends and Learning to Identify High-Risk States and TreatmentsMehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh GhassemiNeurIPS 2021 · 51 citations
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPsHarsh Satija, Philip S. Thomas, Joelle Pineau, Romain LarocheNeurIPS 2021 · 30 citations
- Counterfactual-Augmented Importance Sampling for Semi-Offline Policy EvaluationShengpu Tang, Jenna WiensNeurIPS 2023 · 8 citations
- User-Interactive Offline Reinforcement LearningPhillip Swazinna, Steffen Udluft, Thomas A. RunklerICLR 2023 · 3 citations
Related papers
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 7 citations
- Finding Counterfactually Optimal Action Sequences in Continuous State SpacesStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2023 · 18 citations
- Quasi-optimal Reinforcement Learning with Continuous ActionsYuhan Li, Wenzhuo Zhou, Ruoqing ZhuICLR 2023 · 2 citations
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 6 citations
- Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao Luo, Ke LiICML 2026 · 3 citations
