Beyond Average Value Function in Precision Medicine: Maximum Probability-Driven Reinforcement Learning for Survival Analysis
Jianqi Feng, Chengchun Shi, Zhenke Wu, Xiaodong Yan, Wei Zhao
摘要
Constructing multistage optimal decisions for alternating recurrent event data is critically important in medical and healthcare research. Current reinforcement learning (RL) algorithms have only been applied to time-to-event data, with the objective of maximizing expected survival time. However, alternating recurrent event data has a different structure, which motivates us to model the probability and frequency of event occurrences rather than a single terminal outcome. In this paper, we introduce an RL framework specifically designed for alternating recurrent event data. Our goal is to maximize the probability that the duration between consecutive events exceeds a clinically meaningful threshold. To achieve this, we identify a lower bound of this probability, which transforms the problem into maximizing a cumulative sum of log probabilities, thus enabling direct application of standard RL algorithms. We establish the theoretical properties of the resulting optimal policy and demonstrate through numerical experiments that our proposed algorithm yields a larger probability of that the time between events exceeds a critical threshold compared with existing state-of-the-art algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point ProcessesChao Qu, Xiaoyu Tan, Siqiao Xue, Xiaoming Shi 等AAAI 2023 · 被引用 23 次
- Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued PoliciesShengpu Tang, Aditya Modi, Michael W. Sjoding, Jenna WiensICML 2020 · 被引用 35 次
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 被引用 12 次
- When to Intervene: Learning Optimal Intervention Policies for Critical EventsNiranjan Damera Venkata, Chiranjib BhattacharyyaNeurIPS 2022 · 被引用 8 次
- Recurrent Halting Chain for Early Multi-label ClassificationThomas Hartvigsen, Cansu Sen, Xiangnan Kong, Elke A. RundensteinerKDD 2020 · 被引用 18 次
