Online Meta-Critic Learning for Off-Policy Actor-Critic Methods
Wei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang, Timothy M. Hospedales
摘要
Off-Policy Actor-Critic (Off-PAC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the actor that trains it to take actions with higher expected return. In this paper, we introduce a novel and flexible meta-critic that observes the learning process and meta-learns an additional loss for the actor that accelerates and improves actor-critic learning. Compared to the vanilla critic, the meta-critic network is explicitly trained to accelerate the learning process; and compared to existing meta-learning algorithms, meta-critic is rapidly learned online for a single task, rather than slowly over a family of tasks. Crucially, our meta-critic framework is designed for off-policy based learners, which currently provide state-of-the-art reinforcement learning sample efficiency. We demonstrate that online meta-critic learning leads to improvements in avariety of continuous control environments when combined with contemporary Off-PAC methods DDPG, TD3 and the state-of-the-art SAC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai 等NeurIPS 2022 · 被引用 16 次
- AutoManager: a Meta-Learning Model for Network Management from Intertwined ForecastsAlan Collet, Antonio Bazco Nogueras, Albert Banchs, Marco FioreINFOCOM 2023 · 被引用 10 次
- Stochastic Regret Guarantees for Online Zeroth- and First-Order Bilevel OptimizationParvin Nazari, Bojian Hou, Davoud Ataee Tarzanagh, Li Shen 等NeurIPS 2025 · 被引用 5 次
- Building a Subspace of Policies for Scalable Continual LearningJean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier 等ICLR 2023 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningMotoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta 等ICML 2025
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 被引用 69 次
- Doubly Robust Off-Policy Actor-Critic: Convergence and OptimalityTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangICML 2021 · 被引用 31 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
- Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic LearningHaque Ishfaq, Guangyuan Wang, Sami Nur Islam, Doina PrecupICLR 2025
