PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling
Utsav Singh, Wesley A. Suttle, Brian M. Sadler, Vinay P. Namboodiri, Amrit S. Bedi
Abstract
In this work, we introduce PIPER: Primitive-Informed Preference-based Hierarchical reinforcement learning via Hindsight Relabeling, a novel approach that leverages preference-based learning to learn a reward model, and subsequently uses this reward model to relabel higher-level replay buffers. Since this reward is unaffected by lower primitive behavior, our relabeling-based approach is able to mitigate non-stationarity, which is common in existing hierarchical approaches, and demonstrates impressive performance across a range of challenging sparse-reward tasks. Since obtaining human feedback is typically impractical, we propose to replace the human-in-the-loop approach with our primitive-in-the-loop approach, which generates feedback using sparse rewards provided by the environment. Moreover, in order to prevent infeasible subgoal prediction and avoid degenerate solutions, we propose primitive-informed regularization that conditions higher-level policies to generate feasible subgoals for lower-level policies. We perform extensive experiments to show that PIPER mitigates non-stationarity in hierarchical reinforcement learning and achieves greater than 50 success rates in challenging, sparse-reward robotic environments, where most other baselines fail to achieve any significant progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- A Distributional Approach to Uncertainty-Aware Preference Alignment Using Offline DemonstrationsSheng Xu, Bo Yue, Hongyuan Zha, Guiliang LiuICLR 2025
- Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel ApproachUtsav Singh, Souradip Chakraborty, Wesley Suttle, Brian M. Sadler et al.ICLR 2026
Builds on3
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- Accelerating Robotic Reinforcement Learning via Parameterized Action PrimitivesMurtaza Dalal, Deepak Pathak, Ruslan SalakhutdinovNeurIPS 2021 · 121 citations
Related papers
- CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriAAAI 2026 · 6 citations
- PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriICLR 2025
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu et al.ICLR 2022 · 28 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
