Monte Carlo Augmented Actor-Critic for Sparse Reward Deep Reinforcement Learning from Suboptimal Demonstrations
Albert Wilcox, Ashwin Balakrishna, Jules Dedieu, Wyame Benslimane, Daniel S. Brown, Ken Goldberg
摘要
Providing densely shaped reward functions for RL algorithms is often exceedingly challenging, motivating the development of RL algorithms that can learn from easier-to-specify sparse reward functions. This sparsity poses new exploration challenges. One common way to address this problem is using demonstrations to provide initial signal about regions of the state space with high rewards. However, prior RL from demonstrations algorithms introduce significant complexity and many hyperparameters, making them hard to implement and tune. We introduce Monte Carlo Augmented Actor Critic (MCAC), a parameter free modification to standard actor-critic algorithms which initializes the replay buffer with demonstrations and computes a modified -value by taking the maximum of the standard temporal distance (TD) target and a Monte Carlo estimate of the reward-to-go. This encourages exploration in the neighborhood of high-performing trajectories by encouraging high -values in corresponding regions of the state space. Experiments across continuous control domains suggest that MCAC can be used to significantly increase learning efficiency across commonly used RL and RL-from-demonstrations algorithms. See https://sites.google.com/view/mcac-rl for code and supplementary material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
- Doubly Mild Generalization for Offline Reinforcement LearningYixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang 等NeurIPS 2024 · 被引用 30 次
- Improving Offline RL by Blending HeuristicsSinong Geng, Aldo Pacchiano, Andrey Kolobov, Ching-An ChengICLR 2024 · 被引用 12 次
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng 等NeurIPS 2025 · 被引用 9 次
- Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesZhouyu He, Peng Qiao, Rongchun Li, Yong Dou 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper2
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
相关 Paper
- Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning UpdatesNicholas Corrado, Josiah P. HannaICLR 2024 · 被引用 6 次
- Hybrid Policy Optimization from Imperfect DemonstrationsHanlin Yang, Chao Yu, Peng Sun, Siji ChenNeurIPS 2023 · 被引用 14 次
- Online Meta-Critic Learning for Off-Policy Actor-Critic MethodsWei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang 等NeurIPS 2020 · 被引用 54 次
- Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-CriticWesley A. Suttle, Amrit S. Bedi, Bhrij Patel, Brian M. Sadler 等ICML 2023 · 被引用 24 次
- Memory Based Trajectory-conditioned Policies for Learning from Sparse RewardsYijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng 等NeurIPS 2020 · 被引用 36 次
