Deep Recurrent Optimal Stopping
Niranjan Damera Venkata, Chiranjib Bhattacharyya
摘要
Deep neural networks (DNNs) have recently emerged as a powerful paradigm for solving Markovian optimal stopping problems. However, a ready extension of DNN-based methods to non-Markovian settings requires significant state and parameter space expansion, manifesting the curse of dimensionality. Further, efficient state-space transformations permitting Markovian approximations, such as those afforded by recurrent neural networks (RNNs), are either structurally infeasible or are confounded by the curse of non-Markovianity. Considering these issues, we introduce, for the first time, an optimal stopping policy gradient algorithm (OSPG) that can leverage RNNs effectively in non-Markovian settings by implicitly optimizing value functions without recursion, mitigating the curse of non-Markovianity. The OSPG algorithm is derived from an inference procedure on a novel Bayesian network representation of discrete-time non-Markovian optimal stopping trajectories and, as a consequence, yields an offline policy gradient algorithm that eliminates expensive Monte Carlo policy rollouts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Deep Bayesian Quadrature Policy OptimizationRavi Tej Akella, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Animashree Anandkumar 等AAAI 2021 · 被引用 5 次
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 被引用 19 次
- Structure Learning-Based Task Decomposition for Reinforcement Learning in Non-stationary EnvironmentsHonguk Woo, Gwangpyo Yoo, Minjong YooAAAI 2022 · 被引用 5 次
- DDPNOpt: Differential Dynamic Programming Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICLR 2021 · 被引用 7 次
- Real-Time Recurrent Reinforcement LearningJulian Lemmel, Radu GrosuAAAI 2025 · 被引用 8 次
