Flattening Hierarchies with Policy Bootstrapping
John L. Zhou, Jonathan C. Kao
摘要
Offline goal-conditioned reinforcement learning (GCRL) is a promising approach for pretraining generalist policies on large datasets of reward-free trajectories, akin to the self-supervised objectives used to train foundation models for computer vision and natural language processing. However, scaling GCRL to longer horizons remains challenging due to the combination of sparse rewards and discounting, which obscures the comparative advantages of primitive actions with respect to distant goals. Hierarchical RL methods achieve strong empirical results on long-horizon goal-reaching tasks, but their reliance on modular, timescale-specific policies and subgoal generation introduces significant additional complexity and hinders scaling to high-dimensional goal spaces. In this work, we introduce an algorithm to train a flat (non-hierarchical) goal-conditioned policy by bootstrapping on subgoal-conditioned policies with advantage-weighted importance sampling. Our approach eliminates the need for a generative model over the (sub)goal space, which we find is key for scaling to high-dimensional control in large state spaces. We further show that existing hierarchical and bootstrapping-based approaches correspond to specific design choices within our derivation. Across a comprehensive suite of state- and pixel-based locomotion and manipulation benchmarks, our method matches or surpasses state-of-the-art offline GCRL algorithms and scales to complex, long-horizon tasks where prior approaches fail. Project page: https://johnlyzhou.github.io/saw/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Test-time Offline Reinforcement Learning on Goal-related ExperienceMarco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius 等ICML 2026 · 被引用 7 次
- Test-Time Graph Search for Goal-Conditioned Reinforcement LearningEvgenii Opryshko, Junwei Quan, Claas Voelcker, Yilun Du 等ICML 2026 · 被引用 6 次
- Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement LearningAravind Venugopal, Jiayu Chen, Xudong Wu, Chongyi Zheng 等ICLR 2026 · 被引用 2 次
- Laplacian Representations for Decision-Time PlanningDikshant Shehmar, Matthew Schlegel, Matthew Taylor, Marlos C. MachadoICML 2026 · 被引用 1 次
- QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RLXing Lei, Jincheng Wang, Xuetao Zhang, Donglin WangICML 2026
它引用的顶会 Paper26
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
相关 Paper
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal DiffusionDan Haramati, Carl Qi, Tal Daniel, Amy Zhang 等ICLR 2026 · 被引用 7 次
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 被引用 22 次
- Action-Sufficient Goal RepresentationsJinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup MoonICML 2026 · 被引用 1 次
