Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
Rui Hu, Yu Chen, Longbo Huang
Abstract
Many popular practical reinforcement learning (RL) algorithms employ evolving reward functions—through techniques such as reward shaping, entropy regularization, or curriculum learning—yet their theoretical foundations remain underdeveloped. This paper provides the first finite-time convergence analysis of a single-timescale actor-critic algorithm in the presence of an evolving reward function under Markovian sampling. We consider a setting where the reward parameters may change at each time step, affecting both policy optimization and value estimation. Under standard assumptions, we derive non-asymptotic bounds for both actor and critic errors. Our result shows that an convergence rate is achievable, matching the best-known rate for static rewards, provided the reward parameters evolve slowly enough. This rate is preserved when the reward is updated via a gradient-based rule with bounded gradient and on the same timescale as the actor and critic, offering a theoretical foundation for many popular RL techniques. As a secondary contribution, we introduce a novel analysis of distribution mismatch under Markovian sampling, improving the best-known rate by a factor of in the static-reward case.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e07129d6-ab52-43dc-a972-03d8366c869dBuilds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
Related papers
- Finite-Time Analysis of Single-Timescale Actor-CriticXuyang Chen, Lin ZhaoNeurIPS 2023 · 36 citations
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 189 citations
- Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network ApproximationXuyang Chen, Fengzhuo Zhang, Keyu Yan, Lin ZhaoICLR 2026
- Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function ApproximationYue Wang, Shaofeng Zou, Yi ZhouNeurIPS 2021 · 12 citations
- Non-Asymptotic Analysis for Single-Loop (Natural) Actor-Critic with Compatible Function ApproximationYudan Wang, Yue Wang, Yi Zhou, Shaofeng ZouICML 2024 · 11 citations
