Contrastive UCB: Provably Efficient Contrastive Self-Supervised Learning in Online Reinforcement Learning
Shuang Qiu, Lingxiao Wang, Chenjia Bai, Zhuoran Yang, Zhaoran Wang
摘要
In view of its power in extracting feature representation, contrastive self-supervised learning has been successfully integrated into the practice of (deep) reinforcement learning (RL), leading to efficient policy learning in various applications. Despite its tremendous empirical successes, the understanding of contrastive learning for RL remains elusive. To narrow such a gap, we study how RL can be empowered by contrastive learning in a class of Markov decision processes (MDPs) and Markov games (MGs) with low-rank transitions. For both models, we propose to extract the correct feature representations of the low-rank model by minimizing a contrastive loss. Moreover, under the online setting, we propose novel upper confidence bound (UCB)-type algorithms that incorporate such a contrastive loss with online RL algorithms for MDPs or MGs. We further theoretically prove that our algorithm recovers the true representations and simultaneously achieves sample efficiency in learning the optimal policy and Nash equilibrium in MDPs and MGs. We also provide empirical studies to demonstrate the efficacy of the UCB-based contrastive learning method for RL. To the best of our knowledge, we provide the first provably efficient online RL algorithm that incorporates contrastive learning for representation learning. Our codes are available at https://github.com/Baichenjia/Contrastive-UCB .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 被引用 176 次
- Making Linear MDPs Practical via Contrastive Representation LearningTianjun Zhang, Tongzheng Ren, Mengjiao Yang, Joseph Gonzalez 等ICML 2022 · 被引用 57 次
- Feedback Efficient Online Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali 等ICML 2024 · 被引用 47 次
- Behavior Contrastive Learning for Unsupervised Skill DiscoveryRushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 等ICML 2023 · 被引用 34 次
它引用的顶会 Paper10
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
相关 Paper
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 被引用 138 次
- Learning Markov State Abstractions for Deep Reinforcement LearningCameron Allen, Neev Parikh, Omer Gottesman, George KonidarisNeurIPS 2021 · 被引用 66 次
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang 等AAAI 2023 · 被引用 19 次
- Contrastive Representation for Interactive RecommendationJingyu Li, Zhiyong Feng, Dongxiao He, Hongqi Chen 等AAAI 2025 · 被引用 2 次
- Provable Representation Learning for Imitation with Contrastive Fourier FeaturesOfir Nachum, Mengjiao YangNeurIPS 2021 · 被引用 41 次
