Universal Approximation Theorem of Deep Q-Networks
Qian Qi
摘要
We establish a continuous-time framework for analyzing Deep Q-Networks (DQNs) via stochastic control and Forward-Backward Stochastic Differential Equations (FBSDEs). Considering a continuous-time Markov Decision Process (MDP) driven by a square-integrable martingale, we analyze DQN approximation properties. We show that DQNs can approximate the optimal Q-function on compact sets with arbitrary accuracy and high probability, leveraging residual network approximation theorems and large deviation bounds for the state-action process. We then analyze the convergence of a general Q-learning algorithm for training DQNs in this setting, adapting stochastic approximation theorems. Our analysis emphasizes the interplay between DQN layer count, time discretization, and the role of viscosity solutions (primarily for the value function V * ) in addressing potential non-smoothness of the optimal Q-function. This work bridges deep reinforcement learning and stochastic control, offering insights into DQNs in continuous-time settings, relevant for applications with physical systems or high-frequency data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous EnvironmentsSaket Tiwari, Tejas Kotwal, George Dimitri KonidarisICLR 2026
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- Deep learning for continuous-time stochastic control with jumpsPatrick Cheridito, Jean-Loup Dupret, Donatien HainautNeurIPS 2025 · 被引用 8 次
- Finite-Time Analysis of Adaptive Temporal Difference Learning with Deep Neural NetworksTao Sun, Dongsheng Li, Bao WangNeurIPS 2022 · 被引用 11 次
- Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network ApproximationXuyang Chen, Fengzhuo Zhang, Keyu Yan, Lin ZhaoICLR 2026
