Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL
Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim
摘要
Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient Agent), a unified framework for robust unsupervised reinforcement learning. First, we introduce a skill relabeling mechanism to mitigate non-stationarity and significantly improve data efficiency for pre-training. Second, we propose a Complementary Information Bottleneck (CIB), encouraging the learned skill policy to focus on ego-centric features and become robust to distribution shifts for downstream tasks. Through various experiments, we demonstrate that GenDa significantly enhances the scalability of URL with superior generalizability and data efficiency. Our code and videos are available at https://ihatebroccoli. github.io/official-GenDa/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 被引用 258 次
相关 Paper
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Unsupervised Skill Discovery with Bottleneck Option LearningJaekyeom Kim, Seohong Park, Gunhee KimICML 2021 · 被引用 39 次
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement LearningJinxin Liu, Hao Shen, Donglin Wang, Yachen Kang 等NeurIPS 2021 · 被引用 20 次
- Wasserstein Unsupervised Reinforcement LearningShuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao 等AAAI 2022 · 被引用 30 次
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic SkillsYevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao 等ICML 2021 · 被引用 173 次
