Steady State Analysis of Episodic Reinforcement Learning
Bojun Huang
摘要
This paper proves that the episodic learning environment of every finite-horizon decision task has a unique steady state under any behavior policy, and that the marginal distribution of the agent's input indeed approaches to the steady-state distribution in essentially all episodic learning processes. This observation supports an interestingly reversed mindset against conventional wisdom: While steady states are usually presumed to exist in continual learning and are considered less relevant in episodic learning, it turns out they are guaranteed to exist for the latter. Based on this insight, the paper further develops connections between episodic and continual RL for several important concepts that have been separately treated in the two RL formalisms. Practically, the existence of unique and approachable steady state enables a general, reliable, and efficient way to collect data in episodic RL tasks, which the paper applies to policy gradient algorithms as a demonstration, based on a new steady-state policy gradient theorem. The paper also proposes and empirically evaluates a perturbation method that facilitates rapid mixing in real-world tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Distillation of RL Policies with Formal Guarantees via Variational Abstraction of Markov Decision ProcessesFlorent Delgrange, Ann Nowé, Guillermo A. PérezAAAI 2022 · 被引用 14 次
- Learning to Reach Goals via DiffusionVineet Jain, Siamak RavanbakhshICML 2024 · 被引用 11 次
- The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsRaphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez 等ICLR 2024 · 被引用 10 次
- Langevin Thompson Sampling with Logarithmic Communication: Bandits and Reinforcement LearningAmin Karbasi, Nikki Lijing Kuang, Yi-An Ma, Siddharth MitraICML 2023 · 被引用 7 次
- Balancing Context Length and Mixing Times for Reinforcement Learning at ScaleMatthew Riemer, Khimya Khetarpal, Janarthanan Rajendran, Sarath ChandarNeurIPS 2024 · 被引用 7 次
相关 Paper
- Settling the Horizon-Dependence of Sample Complexity in Reinforcement LearningYuanzhi Li, Ruosong Wang, Lin F. YangFOCS 2021 · 被引用 3 次
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans 等NeurIPS 2025 · 被引用 2 次
- Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient MethodsXin Guo, Anran Hu, Junzi ZhangAAAI 2022 · 被引用 10 次
- Reinforcement Learning in Reward-Mixing MDPsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 23 次
- Continual Learning In Environments With Polynomial Mixing TimesMatthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj 等NeurIPS 2022 · 被引用 18 次
