Quasi-optimal Reinforcement Learning with Continuous Actions
Yuhan Li, Wenzhuo Zhou, Ruoqing Zhu
Abstract
Many real-world applications of reinforcement learning (RL) require making decisions in continuous action environments. In particular, determining the optimal dose level plays a vital role in developing medical treatment regimes. One challenge in adapting existing RL algorithms to medical applications, however, is that the popular infinite support stochastic policies, e.g., Gaussian policy, may assign riskily high dosages and harm patients seriously. Hence, it is important to induce a policy class whose support only contains near-optimal actions, and shrink the action-searching area for effectiveness and reliability. To achieve this, we develop a novel quasi-optimal learning algorithm, which can be easily optimized in off-policy settings with guaranteed convergence under general function approximations. Theoretically, we analyze the consistency, sample complexity, adaptability, and convergence of the proposed algorithm. We evaluate our algorithm with comprehensive simulated experiments and a dose suggestion real application to Ohio Type 1 diabetes dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77afd25a-636a-4bbb-a6ea-e013dad003c9Cited by top-tier papers5
- Bi-Level Offline Policy Optimization with Limited ExplorationWenzhuo ZhouNeurIPS 2023 · 6 citations
- Truncated Gaussian Policy for Debiased Continuous ControlGanghun Lee, Minji Kim, Minsu Lee, Byoung-Tak ZhangAAAI 2025 · 1 citation
- q-exponential family for policy optimizationLingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai et al.ICLR 2025
- Toward Conservative Planning from Human-AI Preferences in Reinforcement LearningHuazhong Wang, Wenzhuo ZhouICLR 2026
- Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse PoliciesLingwei Zhu, Han Wang, Yukie NagaiICLR 2025
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Sparse and Continuous Attention MechanismsAndré F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae et al.NeurIPS 2020 · 55 citations
- Medical Dead-ends and Learning to Identify High-Risk States and TreatmentsMehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh GhassemiNeurIPS 2021 · 51 citations
- Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued PoliciesShengpu Tang, Aditya Modi, Michael W. Sjoding, Jenna WiensICML 2020 · 35 citations
Related papers
- ESCADA: Efficient Safety and Context Aware Dose Allocation for Precision MedicineIlker Demirel, Ahmet Alparslan Celik, Cem TekinNeurIPS 2022 · 6 citations
- Kernel Assisted Learning for Personalized Dose FindingLiangyu Zhu, Wenbin Lu, Michael R. Kosorok, Rui SongKDD 2020 · 4 citations
- Learning to search efficiently for causally near-optimal treatmentsSamuel Håkansson, Viktor Lindblom, Omer Gottesman, Fredrik D. JohanssonNeurIPS 2020 · 7 citations
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesRunze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros et al.NeurIPS 2025 · 7 citations
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 22 citations
