Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point Processes
Chao Qu, Xiaoyu Tan, Siqiao Xue, Xiaoming Shi, James Zhang, Hongyuan Mei
Abstract
We consider a sequential decision making problem where the agent faces the environment characterized by the stochastic discrete events and seeks an optimal intervention policy such that its long-term reward is maximized. This problem exists ubiquitously in social media, finance and health informatics but is rarely investigated by the conventional research in reinforcement learning. To this end, we present a novel framework of the model-based reinforcement learning where the agent's actions and observations are asynchronous stochastic discrete events occurring in continuous-time. We model the dynamics of the environment by Hawkes process with external intervention control term and develop an algorithm to embed such process in the Bellman equation which guides the direction of the value gradient. We demonstrate the superiority of our method in both synthetic simulator and real-data experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- EasyTPP: Towards Open Benchmarking Temporal Point ProcessesSiqiao Xue, Xiaoming Shi, Zhixuan Chu, Yan Wang et al.ICLR 2024 · 53 citations
- Prompt-augmented Temporal Point Process for Streaming Event SequenceSiqiao Xue, Yan Wang, Zhixuan Chu, Xiaoming Shi et al.NeurIPS 2023 · 33 citations
- SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on HierarchiesFan Zhou, Chen Pan, Lintao Ma, Yu Liu et al.AAAI 2023 · 8 citations
- ExPERT: Modeling Human Behavior Under External Stimuli Aware Personalized MTPPSubhendu Khatuya, Ritvik Vij, Paramita Koley, Samik Datta et al.AAAI 2025 · 2 citations
- Amortized Network Intervention to Steer the Excitatory Point ProcessesZitao Song, Wendi Ren, Shuang LiICLR 2024 · 1 citation
Builds on5
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event SequencesSiqiao Xue, Xiaoming Shi, James Y. Zhang, Hongyuan MeiNeurIPS 2022 · 65 citations
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEsJianzhun Du, Joseph Futoma, Finale Doshi-VelezNeurIPS 2020 · 63 citations
- Neural Spatio-Temporal Point ProcessesRicky T. Q. Chen, Brandon Amos, Maximilian NickelICLR 2021 · 21 citations
Related papers
- TREND: TempoRal Event and Node Dynamics for Graph Representation LearningZhihao Wen, Yuan FangWWW 2022 · 113 citations
- POMDPs in Continuous Time and Discrete SpacesBastian Alt, Matthias Schultheis, Heinz KoepplNeurIPS 2020 · 10 citations
- Beyond Average Value Function in Precision Medicine: Maximum Probability-Driven Reinforcement Learning for Survival AnalysisJianqi Feng, Chengchun Shi, Zhenke Wu, Xiaodong Yan et al.NeurIPS 2025 · 1 citation
- Stimuli-Sensitive Hawkes Processes for Personalized Student Procrastination ModelingMengfan Yao, Siqian Zhao, Shaghayegh Sahebi, Reza Feyzi-BehnaghWWW 2021 · 16 citations
- When to Intervene: Learning Optimal Intervention Policies for Critical EventsNiranjan Damera Venkata, Chiranjib BhattacharyyaNeurIPS 2022 · 8 citations
