Simplified Temporal Consistency Reinforcement Learning
Yi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala, Joni Pajarinen
摘要
Reinforcement learning is able to solve complex sequential decision-making tasks but is currently limited by sample efficiency and required computation. To improve sample efficiency, recent work focuses on model-based RL which interleaves model learning with planning. Recent methods further utilize policy learning, value estimation, and, self-supervised learning as auxiliary objectives. In this paper we show that, surprisingly, a simple representation learning approach relying only on a latent dynamics model trained by latent temporal consistency is sufficient for high-performance RL. This applies when using pure planning with a dynamics model conditioned on the representation, but, also when utilizing the representation as policy and value function features in model-free RL. In experiments, our approach learns an accurate dynamics model to solve challenging high-dimensional locomotion tasks with online planners while being 4.1 times faster to train compared to ensemble-based methods. With model-free RL without planning, especially on high-dimensional tasks, such as the DeepMind Control Suite Humanoid and Dog tasks, our approach outperforms model-free methods by a large margin and matches model-based methods' sample efficiency while training 2.4 times faster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma 等ICLR 2024 · 被引用 50 次
- Learning Latent Dynamic Robust Representations for World ModelsRuixiang Sun, Hongyu Zang, Xin Li, Riashat IslamICML 2024 · 被引用 15 次
- COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpacePengyi Li, Hongyao Tang, Yifu Yuan, Jianye Hao 等NeurIPS 2025 · 被引用 3 次
- Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free ControlDevdhar Patel, Hava T. SiegelmannICLR 2025
- Latent Action Learning Requires Supervision in the Presence of DistractorsAlexander Nikulin, Ilya Zisman, Denis Tarasov, Nikita Lyubaykin 等ICML 2025
它引用的顶会 Paper22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
相关 Paper
- Dynamics-Aware EmbeddingsWilliam F. Whitney, Rajat Agarwal, Kyunghyun Cho, Abhinav GuptaICLR 2020
- Simplifying Model-based RL: Learning Representations, Latent-space Models, and Policies with One ObjectiveRaj Ghugare, Homanga Bharadhwaj, Benjamin Eysenbach, Sergey Levine 等ICLR 2023 · 被引用 2 次
- Reward-free World Models for Online Imitation LearningShangzhe Li, Zhiao Huang, Hao SuICML 2025
- TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement LearningRuijie Zheng, Xiyao Wang, Yanchao Sun, Shuang Ma 等NeurIPS 2023 · 被引用 89 次
- Model-Based Reinforcement Learning via Imagination with Derived MemoryYao Mu, Yuzheng Zhuang, Bin Wang, Guangxiang Zhu 等NeurIPS 2021 · 被引用 14 次
