Warm-Start Actor-Critic: From Approximation Error to Sub-optimality Gap
Hang Wang, Sen Lin, Junshan Zhang
Abstract
Warm-Start reinforcement learning (RL), aided by a prior policy obtained from offline training, is emerging as a promising RL approach for practical applications. Recent empirical studies have demonstrated that the performance of Warm-Start RL can be improved quickly in some cases but become stagnant in other cases, especially when the function approximation is used. To this end, the primary objective of this work is to build a fundamental understanding on ``whether and when online learning can be significantly accelerated by a warm-start policy from offline RL?''. Specifically, we consider the widely used Actor-Critic (A-C) method with a prior policy. We first quantify the approximation errors in the Actor update and the Critic update, respectively. Next, we cast the Warm-Start A-C algorithm as Newton's method with perturbation, and study the impact of the approximation errors on the finite-time learning performance with inaccurate Actor/Critic updates. Under some general technical conditions, we derive the upper bounds, which shed light on achieving the desired finite-learning performance in the Warm-Start A-C algorithm. In particular, our findings reveal that it is essential to reduce the algorithm bias in online learning. We also obtain lower bounds on the sub-optimality gap of the Warm-Start A-C algorithm to quantify the impact of the bias and error propagation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 110 citations
- A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time BoundGal Dalal, Balázs Szörényi, Gugan ThoppeAAAI 2020 · 59 citations
- Single-Timescale Actor-Critic Provably Finds Globally Optimal PolicyZuyue Fu, Zhuoran Yang, Zhaoran WangICLR 2021 · 52 citations
Related papers
- Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline DataZhiyuan Zhou, Andy Peng, Qiyang Li, Sergey Levine et al.ICLR 2025
- Actor-Critics Can Achieve Optimal Sample EfficiencyKevin Tan, Wei Fan, Yuting WeiICML 2025
- A Perspective of Q-value Estimation on Offline-to-Online Reinforcement LearningYinmin Zhang, Jie Liu, Chuming Li, Yazhe Niu et al.AAAI 2024 · 28 citations
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement LearningAndrea Zanette, Martin J. Wainwright, Emma BrunskillNeurIPS 2021 · 140 citations
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 87 citations
