Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research
Michal Bortkiewicz, Wladyslaw Palucki, Vivek Myers, Tadeusz Dziarmaga, Tomasz Arczewski, Lukasz Kucinski, Benjamin Eysenbach
摘要
Self-supervision has the potential to transform reinforcement learning (RL), paralleling the breakthroughs it has enabled in other areas of machine learning. While self-supervised learning in other domains aims to find patterns in a fixed dataset, self-supervised goal-conditioneds reinforcement learning (GCRL) agents discover new behaviors by learning from the goals achieved during unstructured interaction with the environment. However, these methods have failed to see similar success, both due to a lack of data from slow environment simulations as well as a lack of stable algorithms. We take a step toward addressing both of these issues by releasing a high-performance codebase and benchmark (JaxGCRL) for self-supervised GCRL, enabling researchers to train agents for millions of environment steps in minutes on a single GPU. By utilizing GPU-accelerated replay buffers, environments, and a stable contrastive RL algorithm, we reduce training time by up to 22×. Additionally, we assess key design choices in contrastive RL, identifying those that most effectively stabilize and enhance training performance. With this approach, we provide a foundation for future research in self-supervised GCRL, enabling researchers to quickly iterate on new ideas and evaluate them in diverse and challenging environments. Code: https://github.com/MichalBortkiewicz/JaxGCRL.
To achieve this training and performance improvement, we combine insights from self-supervised RL with recent advances in GPU-accelerated simulation. The first key ingredient is recent work on GPU-accelerated simulators, both for physics (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Demystifying The Mechanisms Behind Emergent Exploration in Goal-Conditioned RLMahsa Bastankhah, Grace Liu, Dilip Arumugam, Thomas L. Griffiths 等ICLR 2026 · 被引用 7 次
- Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic RewardsFaisal Mohamed, Catherine Ji, Benjamin Eysenbach, Glen BersethICLR 2026 · 被引用 1 次
- QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RLXing Lei, Jincheng Wang, Xuetao Zhang, Donglin WangICML 2026
- Normalizing Flows are Capable Models for Continuous ControlRaj Ghugare, Benjamin EysenbachNeurIPS 2025
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 被引用 347 次
相关 Paper
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching CapabilitiesKevin Wang, Ishaan Javali, Michal Bortkiewicz, Tomasz Trzcinski 等NeurIPS 2025 · 被引用 46 次
- Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline DataChongyi Zheng, Benjamin Eysenbach, Homer Rich Walke, Patrick Yin 等ICLR 2024 · 被引用 16 次
- Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control TasksMichael T. Matthews, Michael Beukman, Chris Lu, Jakob Nicolaus FoersterICLR 2025
- Rethinking Goal-Conditioned Supervised Learning and Its Connection to Offline RLRui Yang, Yiming Lu, Wenzhe Li, Hao Sun 等ICLR 2022 · 被引用 100 次
- Provably Efficient Offline Goal-Conditioned Reinforcement Learning with General Function Approximation and Single-Policy ConcentrabilityHanlin Zhu, Amy ZhangNeurIPS 2023 · 被引用 7 次
