TAAC: Temporally Abstract Actor-Critic for Continuous Control
Haonan Yu, Wei Xu, Haichao Zhang
摘要
We present temporally abstract actor-critic (TAAC), a simple but effective offpolicy RL algorithm that incorporates closed-loop temporal abstraction into the actor-critic framework. TAAC adds a second-stage binary policy to choose between the previous action and a new action output by an actor. Crucially, its "act-or-repeat" decision hinges on the actually sampled action instead of the expected behavior of the actor. This post-acting switching scheme let the overall policy make more informed decisions. TAAC has two important features: a) persistent exploration, and b) a new compare-through Q operator for multi-step TD backup, specially tailored to the action repetition scenario. We demonstrate TAAC's advantages over several strong baselines across 14 continuous control tasks. Our surprising finding reveals that while achieving top performance, TAAC is able to "mine" a significant number of repeated actions with the trained policy even on continuous tasks whose problem structures on the surface seem to repel action repetition. This suggests that aside from encouraging persistent exploration, action repetition can find its place in a good policy behavior. Code is available at https://github.com/hnyu/taac . 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li 等ICML 2023 · 被引用 19 次
- Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement LearningGe Li, Hongyi Zhou, Dominik Roth, Serge Thilges 等ICLR 2024 · 被引用 11 次
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler 等NeurIPS 2024 · 被引用 10 次
- Policy Expansion for Bridging Offline-to-Online Reinforcement LearningHaichao Zhang, Wei Xu, Haonan YuICLR 2023 · 被引用 5 次
- LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option FrameworkWoojun Kim, Jeonghye Kim, Youngchul SungICML 2023 · 被引用 5 次
它引用的顶会 Paper9
- Learning to Coordinate Manipulation Skills via Skill Behavior DiversificationYoungwoon Lee, Jingyun Yang, Joseph J. LimICLR 2020 · 被引用 98 次
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 被引用 97 次
- Sub-policy Adaptation for Hierarchical Reinforcement LearningAlexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter AbbeelICLR 2020 · 被引用 85 次
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni 等ICML 2020 · 被引用 43 次
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 被引用 38 次
相关 Paper
- Learning Uncertainty-Aware Temporally-Extended ActionsJoongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan OhAAAI 2024 · 被引用 3 次
- Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features CriticsLuca Grillotti, Maxence Faldor, Borja G. León, Antoine CullyICML 2024 · 被引用 13 次
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 被引用 38 次
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
- Online Meta-Critic Learning for Off-Policy Actor-Critic MethodsWei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang 等NeurIPS 2020 · 被引用 54 次
