Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
Ben Eysenbach, Xinyang Geng, Sergey Levine, Ruslan Salakhutdinov
摘要
Multi-task reinforcement learning (RL) aims to simultaneously learn policies for solving many tasks. Several prior works have found that relabeling past experience with different reward functions can improve sample efficiency. Relabeling methods typically ask: if, in hindsight, we assume that our experience was optimal for some task, for what task was it optimal? In this paper, we show that hindsight relabeling is inverse RL, an observation that suggests that we can use inverse RL in tandem for RL algorithms to efficiently solve many tasks. We use this idea to generalize goal-relabeling techniques from prior work to arbitrary classes of tasks. Our experiments confirm that relabeling data using inverse RL accelerates learning in general multi-task settings, including goal-reaching, domains with discrete sets of rewards, and those with linear reward functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 331 次
- RvS: What is Essential for Offline RL via Supervised Learning?Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, Sergey LevineICLR 2022 · 被引用 225 次
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic SkillsYevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao 等ICML 2021 · 被引用 173 次
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 被引用 125 次
- Rethinking Goal-Conditioned Supervised Learning and Its Connection to Offline RLRui Yang, Yiming Lu, Wenzhe Li, Hao Sun 等ICLR 2022 · 被引用 100 次
它引用的顶会 Paper1
相关 Paper
- Generalized Hindsight for Reinforcement LearningAlexander C. Li, Lerrel Pinto, Pieter AbbeelNeurIPS 2020 · 被引用 81 次
- Hindsight Foresight Relabeling for Meta-Reinforcement LearningMichael Wan, Jian Peng, Tanmay GangwaniICLR 2022 · 被引用 7 次
- Hindsight Task Relabelling: Experience Replay for Sparse Reward Meta-RLCharles Packer, Pieter Abbeel, Joseph E. GonzalezNeurIPS 2021 · 被引用 22 次
- How Does Goal Relabeling Improve Sample Efficiency?Sirui Zheng, Chenjia Bai, Zhuoran Yang, Zhaoran WangICML 2024 · 被引用 5 次
- Correcting experience replay for multi-agent communicationSanjeevan Ahilan, Peter DayanICLR 2021 · 被引用 3 次
