A State-Distribution Matching Approach to Non-Episodic Reinforcement Learning
Archit Sharma, Rehaan Ahmad, Chelsea Finn
Abstract
While reinforcement learning (RL) provides a framework for learning through trial and error, translating RL algorithms into the real world has remained challenging. A major hurdle to realworld application arises from the development of algorithms in an episodic setting where the environment is reset after every trial, in contrast with the continual and non-episodic nature of the realworld encountered by embodied agents such as humans and robots. Enabling agents to learn behaviors autonomously in such non-episodic environments requires that the agent to be able to conduct its own trials. Prior works have considered an alternating approach where a forward policy learns to solve the task and the backward policy learns to reset the environment, but what initial state distribution should the backward policy reset the agent to? Assuming access to a few demonstrations, we propose a new method, MEDAL, that trains the backward policy to match the state distribution in the provided demonstrations. This keeps the agent close to the task-relevant states, allowing for a mix of easy and difficult starting states for the forward policy. Our experiments show that MEDAL matches or outperforms prior methods on three sparse-reward continuous control tasks from the EARL benchmark, with 40% gains on the hardest task, while making fewer assumptions than prior works. Code and videos are at: https://sites.google.com/view/medal-arl/home
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acf6904b-d872-40d1-bb9d-fb7f04396309Cited by top-tier papers9
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 33 citations
- When to Ask for Help: Proactive Interventions in Autonomous Reinforcement LearningAnnie Xie, Fahim Tajwar, Archit Sharma, Chelsea FinnNeurIPS 2022 · 30 citations
- NeoRL: Efficient Exploration for Nonepisodic RLBhavya Sukhija, Lenart Treven, Florian Dörfler, Stelian Coros et al.NeurIPS 2024 · 7 citations
- Hybrid Reinforcement Learning from Offline Observation AloneYuda Song, Drew Bagnell, Aarti SinghICML 2024 · 6 citations
- Demonstration-free Autonomous Reinforcement Learning via Implicit and Bidirectional CurriculumJigang Kim, Daesol Cho, H. Jin KimICML 2023 · 4 citations
Builds on9
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah et al.ICLR 2020 · 202 citations
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher et al.ICML 2020 · 178 citations
- Off-Policy Imitation Learning from ObservationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouNeurIPS 2020 · 102 citations
- Visual Adversarial Imitation Learning using Variational ModelsRafael Rafailov, Tianhe Yu, Aravind Rajeswaran, Chelsea FinnNeurIPS 2021 · 61 citations
Related papers
- Autonomous Reinforcement Learning: Formalism and BenchmarkingArchit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta et al.ICLR 2022 · 39 citations
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil et al.NeurIPS 2022 · 2 citations
- Autonomous Reinforcement Learning via Subgoal CurriculaArchit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman et al.NeurIPS 2021 · 41 citations
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski et al.NeurIPS 2021 · 152 citations
- Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple BaselineHongjoon Ahn, Jinu Hyeon, Youngmin Oh, Bosun Hwang et al.ICLR 2025
