Deep Recurrent Optimal Stopping
Niranjan Damera Venkata, Chiranjib Bhattacharyya
Abstract
Deep neural networks (DNNs) have recently emerged as a powerful paradigm for solving Markovian optimal stopping problems. However, a ready extension of DNN-based methods to non-Markovian settings requires significant state and parameter space expansion, manifesting the curse of dimensionality. Further, efficient state-space transformations permitting Markovian approximations, such as those afforded by recurrent neural networks (RNNs), are either structurally infeasible or are confounded by the curse of non-Markovianity. Considering these issues, we introduce, for the first time, an optimal stopping policy gradient algorithm (OSPG) that can leverage RNNs effectively in non-Markovian settings by implicitly optimizing value functions without recursion, mitigating the curse of non-Markovianity. The OSPG algorithm is derived from an inference procedure on a novel Bayesian network representation of discrete-time non-Markovian optimal stopping trajectories and, as a consequence, yields an offline policy gradient algorithm that eliminates expensive Monte Carlo policy rollouts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Deep Bayesian Quadrature Policy OptimizationRavi Tej Akella, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Animashree Anandkumar et al.AAAI 2021 · 5 citations
- SVQN: Sequential Variational Soft Q-Learning NetworksShiyu Huang, Hang Su, Jun Zhu, Ting ChenICLR 2020 · 19 citations
- Structure Learning-Based Task Decomposition for Reinforcement Learning in Non-stationary EnvironmentsHonguk Woo, Gwangpyo Yoo, Minjong YooAAAI 2022 · 5 citations
- DDPNOpt: Differential Dynamic Programming Neural OptimizerGuan-Horng Liu, Tianrong Chen, Evangelos A. TheodorouICLR 2021 · 7 citations
- Real-Time Recurrent Reinforcement LearningJulian Lemmel, Radu GrosuAAAI 2025 · 8 citations
