Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point Processes
Chao Qu, Xiaoyu Tan, Siqiao Xue, Xiaoming Shi, James Zhang, Hongyuan Mei
摘要
We consider a sequential decision making problem where the agent faces the environment characterized by the stochastic discrete events and seeks an optimal intervention policy such that its long-term reward is maximized. This problem exists ubiquitously in social media, finance and health informatics but is rarely investigated by the conventional research in reinforcement learning. To this end, we present a novel framework of the model-based reinforcement learning where the agent's actions and observations are asynchronous stochastic discrete events occurring in continuous-time. We model the dynamics of the environment by Hawkes process with external intervention control term and develop an algorithm to embed such process in the Bellman equation which guides the direction of the value gradient. We demonstrate the superiority of our method in both synthetic simulator and real-data experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- EasyTPP: Towards Open Benchmarking Temporal Point ProcessesSiqiao Xue, Xiaoming Shi, Zhixuan Chu, Yan Wang 等ICLR 2024 · 被引用 53 次
- Prompt-augmented Temporal Point Process for Streaming Event SequenceSiqiao Xue, Yan Wang, Zhixuan Chu, Xiaoming Shi 等NeurIPS 2023 · 被引用 33 次
- SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on HierarchiesFan Zhou, Chen Pan, Lintao Ma, Yu Liu 等AAAI 2023 · 被引用 8 次
- ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPPSubhendu Khatuya, Ritvik Vij, Paramita Koley, Samik Datta 等AAAI 2025 · 被引用 2 次
- Amortized Network Intervention to Steer the Excitatory Point ProcessesZitao Song, Wendi Ren, Shuang LiICLR 2024 · 被引用 1 次
它引用的顶会 Paper5
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event SequencesSiqiao Xue, Xiaoming Shi, James Y. Zhang, Hongyuan MeiNeurIPS 2022 · 被引用 65 次
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEsJianzhun Du, Joseph Futoma, Finale Doshi-VelezNeurIPS 2020 · 被引用 63 次
- Neural Spatio-Temporal Point ProcessesRicky T. Q. Chen, Brandon Amos, Maximilian NickelICLR 2021 · 被引用 21 次
相关 Paper
- TREND: TempoRal Event and Node Dynamics for Graph Representation LearningZhihao Wen, Yuan FangWWW 2022 · 被引用 113 次
- POMDPs in Continuous Time and Discrete SpacesBastian Alt, Matthias Schultheis, Heinz KoepplNeurIPS 2020 · 被引用 10 次
- Beyond Average Value Function in Precision Medicine: Maximum Probability-Driven Reinforcement Learning for Survival AnalysisJianqi Feng, Chengchun Shi, Zhenke Wu, Xiaodong Yan 等NeurIPS 2025 · 被引用 1 次
- Stimuli-Sensitive Hawkes Processes for Personalized Student Procrastination ModelingMengfan Yao, Siqian Zhao, Shaghayegh Sahebi, Reza Feyzi-BehnaghWWW 2021 · 被引用 16 次
- When to Intervene: Learning Optimal Intervention Policies for Critical EventsNiranjan Damera Venkata, Chiranjib BhattacharyyaNeurIPS 2022 · 被引用 8 次
