Training Transition Policies via Distribution Matching for Complex Tasks
Ju-Seung Byun, Andrew Perrault
摘要
Humans decompose novel complex tasks into simpler ones to exploit previously learned skills. Analogously, hierarchical reinforcement learning seeks to leverage lower-level policies for simple tasks to solve complex ones. However, because each lower-level policy induces a different distribution of states, transitioning from one lower-level policy to another may fail due to an unexpected starting state. We introduce transition policies that smoothly connect lower-level policies by producing a distribution of states and actions that matches what is expected by the next policy. Training transition policies is challenging because the natural reward signal -- whether the next policy can execute its subtask successfully -- is sparse. By training transition policies via adversarial inverse reinforcement learning to match the distribution of expected states and actions, we avoid relying on task-based reward. To further improve performance, we use deep Q-learning with a binary action space to determine when to switch from a transition policy to the next pre-trained policy, using the success or failure of the next subtask as the reward. Although the reward is still sparse, the problem is less severe due to the simple binary action space. We demonstrate our method on continuous bipedal locomotion and arm manipulation tasks that require diverse skills. We show that it smoothly connects the lower-level policies, achieving higher success rates than previous methods that search for successful trajectories based on a reward function, but do not match the state distribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Hierarchical Skills for Efficient ExplorationJonas Gehring, Gabriel Synnaeve, Andreas Krause, Nicolas UsunierNeurIPS 2021 · 被引用 52 次
- Multi-task Hierarchical Adversarial Inverse Reinforcement LearningJiayu Chen, Dipesh Tamboli, Tian Lan, Vaneet AggarwalICML 2023 · 被引用 19 次
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson 等ICLR 2020 · 被引用 35 次
- Skill Discovery for Exploration and Planning using Deep Skill GraphsAkhil Bagaria, Jason K. Senthil, George KonidarisICML 2021 · 被引用 73 次
- State-Conditioned Adversarial Subgoal GenerationVivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian KämäräinenAAAI 2023 · 被引用 16 次
