Hierarchical Reinforcement Learning by Discovering Intrinsic Options
Jesse Zhang, Haonan Yu, Wei Xu
摘要
We propose a hierarchical reinforcement learning method, HIDIO, that can learn task-agnostic options in a self-supervised manner while jointly learning to utilize them to solve sparse-reward tasks. Unlike current hierarchical RL approaches that tend to formulate goal-reaching low-level tasks or pre-define ad hoc lowerlevel policies, HIDIO encourages lower-level option learning that is independent of the task at hand, requiring few assumptions or little knowledge about the task structure. These options are learned through an intrinsic entropy minimization objective conditioned on the option sub-trajectories. The learned options are diverse and task-agnostic. In experiments on sparse-reward robotic manipulation and navigation tasks, HIDIO achieves higher success rates with greater sample efficiency than regular RL baselines and two state-of-the-art hierarchical RL methods. Code available at https://www.github.com/jesbu1/hidio .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Lipschitz-constrained Unsupervised Skill DiscoverySeohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee 等ICLR 2022 · 被引用 72 次
- Hierarchical Skills for Efficient ExplorationJonas Gehring, Gabriel Synnaeve, Andreas Krause, Nicolas UsunierNeurIPS 2021 · 被引用 52 次
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon ReasoningDhruv Shah, Peng Xu, Yao Lu, Ted Xiao 等ICLR 2022 · 被引用 50 次
- HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination MechanismZhiwei Xu, Yunpeng Bai, Bin Zhang, Dapeng Li 等AAAI 2023 · 被引用 46 次
它引用的顶会 Paper6
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 被引用 126 次
- Learning to Coordinate Manipulation Skills via Skill Behavior DiversificationYoungwoon Lee, Jingyun Yang, Joseph J. LimICLR 2020 · 被引用 98 次
- Sub-policy Adaptation for Hierarchical Reinforcement LearningAlexander C. Li, Carlos Florensa, Ignasi Clavera, Pieter AbbeelICLR 2020 · 被引用 85 次
相关 Paper
- Possibility Before Utility: Learning And Using Hierarchical AffordancesRobby Costales, Shariq Iqbal, Fei ShaICLR 2022 · 被引用 5 次
- Hierarchical Planning and Learning for Robots in Stochastic Settings Using Zero-Shot Option InventionNaman Shah, Siddharth SrivastavaAAAI 2024 · 被引用 3 次
- PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriICLR 2025
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu 等ICLR 2022 · 被引用 28 次
- Offline Hierarchical Reinforcement Learning via Inverse OptimizationCarolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone 等ICLR 2025
