Unsupervised Behavior Extraction via Random Intent Priors
Hao Hu, Yiqin Yang, Jianing Ye, Ziqing Mai, Chongjie Zhang
摘要
Reward-free data is abundant and contains rich prior knowledge of human behaviors, but it is not well exploited by offline reinforcement learning (RL) algorithms. In this paper, we propose UBER, an unsupervised approach to extract useful behaviors from offline reward-free datasets via diversified rewards. UBER assigns different pseudo-rewards sampled from a given prior distribution to different agents to extract a diverse set of behaviors, and reuse them as candidate policies to facilitate the learning of new tasks. Perhaps surprisingly, we show that rewards generated from random neural networks are sufficient to extract diverse and useful behaviors, some even close to expert ones. We provide both empirical and theoretical evidence to justify the use of random priors for the reward function. Experiments on multiple benchmarks showcase UBER's ability to learn effective and diverse behavior sets that enhance sample efficiency for online RL, outperforming existing baselines. By reducing reliance on human supervision, UBER broadens the applicability of RL to real-world scenarios with abundant reward-free data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Unsupervised Zero-Shot Reinforcement Learning via Functional Reward EncodingsKevin Frans, Seohong Park, Pieter Abbeel, Sergey LevineICML 2024 · 被引用 26 次
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningAndrew Wagenmaker, Perry Dong, Raymond Tsao, Chelsea Finn 等ICML 2026 · 被引用 10 次
它引用的顶会 Paper34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
相关 Paper
- Leveraging Skills from Unlabeled Prior Data for Efficient Online ExplorationMax Wilcoxson, Qiyang Li, Kevin Frans, Sergey LevineICML 2025
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
- DIDI: Diffusion-Guided Diversity for Offline Behavioral GenerationJinxin Liu, Xinghong Guo, Zifeng Zhuang, Donglin WangICML 2024 · 被引用 3 次
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement LearningAnurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine 等ICLR 2021 · 被引用 15 次
- Future-conditioned Unsupervised Pretraining for Decision TransformerZhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu 等ICML 2023 · 被引用 32 次
