C-Learning: Learning to Achieve Goals via Recursive Classification
Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine
摘要
We study the problem of predicting and controlling the future state distribution of an autonomous agent. This problem, which can be viewed as a reframing of goal-conditioned reinforcement learning (RL), is centered around learning a conditional probability density function over future states. Instead of directly estimating this density function, we indirectly estimate this density function by training a classifier to predict whether an observation comes from the future. Via Bayes' rule, predictions from our classifier can be transformed into predictions over future states. Importantly, an off-policy variant of our algorithm allows us to predict the future state distribution of a new policy, without collecting new experience. This variant allows us to optimize functionals of a policy's future state distribution, such as the density of reaching a particular goal state. While conceptually similar to Q-learning, our work lays a principled foundation for goal-conditioned RL as density estimation, providing justification for goal-conditioned methods used in prior work. This foundation makes hypotheses about Q-learning, including the optimal goal-sampling ratio, which we confirm experimentally. Moreover, our proposed method is competitive with prior goal-conditioned RL methods. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 被引用 173 次
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 被引用 140 次
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 被引用 125 次
- Adversarial Intrinsic Motivation for Reinforcement LearningIshan Durugkar, Mauricio Tec, Scott Niekum, Peter StoneNeurIPS 2021 · 被引用 61 次
它引用的顶会 Paper5
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie 等ICML 2020 · 被引用 145 次
- Rewriting History with Inverse RL: Hindsight Inference for Policy ImprovementBen Eysenbach, Xinyang Geng, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 96 次
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 被引用 71 次
- Neural Topological SLAM for Visual NavigationDevendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, Saurabh GuptaCVPR 2020
相关 Paper
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh 等ICLR 2026 · 被引用 13 次
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching TasksTianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine 等ICLR 2022 · 被引用 19 次
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang 等ICCV 2025 · 被引用 13 次
- Offline Goal-Conditioned Reinforcement Learning via -Advantage RegressionYecheng Jason Ma, Jason Yan, Dinesh Jayaraman, Osbert BastaniNeurIPS 2022 · 被引用 26 次
- Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelJing Zhang, Linjiajie Fang, Kexin Shi, Wenjia Wang 等NeurIPS 2024 · 被引用 14 次
