Offline Goal-Conditioned Reinforcement Learning via -Advantage Regression
Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, Osbert Bastani
摘要
Offline goal-conditioned reinforcement learning (GCRL) promises general-purpose skill learning in the form of reaching diverse goals from purely offline datasets. We propose Goal-conditioned f -Advantage Regression (GoFAR), a novel regressionbased offline GCRL algorithm derived from a state-occupancy matching perspective; the key intuition is that the goal-reaching task can be formulated as a stateoccupancy matching problem between a dynamics-abiding imitator agent and an expert agent that directly teleports to the goal. In contrast to prior approaches, GoFAR does not require any hindsight relabeling and enjoys uninterleaved optimization for its value and policy networks. These distinct features confer GoFAR with much better offline performance and stability as well as statistical performance guarantee that is unattainable for prior methods. Furthermore, we demonstrate that GoFAR's training objectives can be re-purposed to learn an agent-independent goal-conditioned planner from purely offline source-domain data, which enables zero-shot transfer to new target domains. Through extensive experiments, we validate GoFAR's effectiveness in various problem settings and tasks, significantly outperforming prior state-of-art. Notably, on a real robotic dexterous manipulation task, while no other method makes meaningful progress, GoFAR acquires complex manipulation behavior that successfully accomplishes diverse goals. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 被引用 80 次
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 被引用 22 次
- Distance Weighted Supervised Learning for Offline Interaction DataJoey Hejna, Jensen Gao, Dorsa SadighICML 2023 · 被引用 21 次
- DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningJianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao 等ICML 2024 · 被引用 19 次
- Stitching Sub-trajectories with Conditional Diffusion Model for Goal-Conditioned Offline RLSungyoon Kim, Yunseon Choi, Daiki E. Matsunaga, Kee-Eung KimAAAI 2024 · 被引用 19 次
它引用的顶会 Paper14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 被引用 183 次
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic SkillsYevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao 等ICML 2021 · 被引用 173 次
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 被引用 141 次
相关 Paper
- Score Models for Offline Goal-Conditioned Reinforcement LearningHarshit Sikchi, Rohan Chitnis, Ahmed Touati, Alborz Geramifard 等ICLR 2024 · 被引用 16 次
- Provably Efficient Offline Goal-Conditioned Reinforcement Learning with General Function Approximation and Single-Policy ConcentrabilityHanlin Zhu, Amy ZhangNeurIPS 2023 · 被引用 7 次
- What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?Rui Yang, Lin Yong, Xiaoteng Ma, Hao Hu 等ICML 2023 · 被引用 35 次
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 被引用 48 次
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu 等ICML 2025
