A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
Grace Liu, Michael Tang, Benjamin Eysenbach
摘要
In this paper, we present empirical evidence of skills and directed exploration emerging from a simple RL algorithm long before any successful trials are observed. For example, in a manipulation task, the agent is given a single observation of the goal state (see Fig. 1 ) and learns skills, first moving its end-effector, then pushing the block, and finally lifting and placing the block. These skills emerge before the agent has ever successfully placed the block at the goal location and without the aid of any reward functions, demonstrations, or manually-specified distance metrics. Implementing our method involves a simple modification of prior work and does not require density estimates, ensembles, or any additional hyperparameters. We lack a clear theoretical understanding of why the method works so effectively, though our experiments provide some hints. Videos and code: https://graliuce.github.io/sgcrl Sawyer bin Training progress Open/close gripper Push object Roll object between bins Pick and place Move hand left/right move hand front/back Pick up the object Knock object out of bin Figure 1: Skills and Directed Exploration Emerge. In this bin picking task, we provide the agent with a single goal observation where the green block is in the left bin. The agent never receives any rewards. Throughout the course of training, the agent learns skills that increase in complexity. Easier skills seem to enable the agent to unlock more complex skills: moving the hand is a prerequisite for pushing the object; closing the gripper is a prerequisite for picking up the object, which is a prerequisite for moving the object. Figures 7 and 8 in Appendix D show similar progress milestones for other tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 被引用 22 次
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths 等ICLR 2026 · 被引用 7 次
- In-Context Learning for Pure ExplorationAlessio Russo, Ryan Welch, Aldo PacchianoICLR 2026 · 被引用 5 次
- Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement LearningPatrik Reizinger, Bálint Mucsányi, Siyuan Guo, Benjamin Eysenbach 等ICLR 2026 · 被引用 4 次
- Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic RewardsFaisal Mohamed, Catherine Ji, Benjamin Eysenbach, Glen BersethICLR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
相关 Paper
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 被引用 14 次
- Controllability-Aware Unsupervised Skill DiscoverySeohong Park, Kimin Lee, Youngwoon Lee, Pieter AbbeelICML 2023 · 被引用 62 次
- Unsupervised Skill Discovery for Learning Shared Structures across Changing EnvironmentsSang-Hyun Lee, Seung-Woo SeoICML 2023 · 被引用 6 次
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement LearningJinxin Liu, Hao Shen, Donglin Wang, Yachen Kang 等NeurIPS 2021 · 被引用 20 次
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats 等NeurIPS 2022 · 被引用 62 次
