C-Learning: Learning to Achieve Goals via Recursive Classification
Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine
Abstract
We study the problem of predicting and controlling the future state distribution of an autonomous agent. This problem, which can be viewed as a reframing of goal-conditioned reinforcement learning (RL), is centered around learning a conditional probability density function over future states. Instead of directly estimating this density function, we indirectly estimate this density function by training a classifier to predict whether an observation comes from the future. Via Bayes' rule, predictions from our classifier can be transformed into predictions over future states. Importantly, an off-policy variant of our algorithm allows us to predict the future state distribution of a new policy, without collecting new experience. This variant allows us to optimize functionals of a policy's future state distribution, such as the density of reaching a particular goal state. While conceptually similar to Q-learning, our work lays a principled foundation for goal-conditioned RL as density estimation, providing justification for goal-conditioned methods used in prior work. This foundation makes hypotheses about Q-learning, including the optimal goal-sampling ratio, which we confirm experimentally. Moreover, our proposed method is competitive with prior goal-conditioned RL methods. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 125 citations
- Adversarial Intrinsic Motivation for Reinforcement LearningIshan Durugkar, Mauricio Tec, Scott Niekum, Peter StoneNeurIPS 2021 · 61 citations
Builds on5
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Rewriting History with Inverse RL: Hindsight Inference for Policy ImprovementBen Eysenbach, Xinyang Geng, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2020 · 96 citations
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 71 citations
- Neural Topological SLAM for Visual NavigationDevendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, Saurabh GuptaCVPR 2020
Related papers
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh et al.ICLR 2026 · 13 citations
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching TasksTianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine et al.ICLR 2022 · 19 citations
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang et al.ICCV 2025 · 13 citations
- Offline Goal-Conditioned Reinforcement Learning via -Advantage RegressionYecheng Jason Ma, Jason Yan, Dinesh Jayaraman, Osbert BastaniNeurIPS 2022 · 26 citations
- Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelJing Zhang, Linjiajie Fang, Kexin Shi, Wenjia Wang et al.NeurIPS 2024 · 14 citations
