Autonomous Reinforcement Learning: Formalism and Benchmarking
Archit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta, Karol Hausman, Sergey Levine, Chelsea Finn
摘要
Reinforcement learning (RL) provides a naturalistic framing for learning through trial and error, which is appealing both because of its simplicity and effectiveness and because of its resemblance to how humans and animals acquire skills through experience. However, real-world embodied learning, such as that performed by humans and animals, is situated in a continual, non-episodic world, whereas common benchmark tasks in RL are episodic, with the environment resetting between trials to provide the agent with multiple attempts. This discrepancy presents a major challenge when attempting to take RL algorithms developed for episodic simulated environments and run them on real-world platforms, such as robots. In this paper, we aim to address this discrepancy by laying out a framework for Autonomous Reinforcement Learning (ARL): reinforcement learning where the agent not only learns through its own experience, but also contends with lack of human supervision to reset between trials. We introduce a simulated benchmark EARL 1 around this framework, containing a set of diverse and challenging simulated tasks reflective of the hurdles introduced to learning when only a minimal reliance on extrinsic intervention can be assumed. We show that standard approaches to episodic RL and existing approaches struggle as interventions are minimized, underscoring the need for developing new algorithms for reinforcement learning with a greater focus on autonomy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- You Only Live Once: Single-Life Reinforcement LearningAnnie S. Chen, Archit Sharma, Sergey Levine, Chelsea FinnNeurIPS 2022 · 被引用 33 次
- When to Ask for Help: Proactive Interventions in Autonomous Reinforcement LearningAnnie Xie, Fahim Tajwar, Archit Sharma, Chelsea FinnNeurIPS 2022 · 被引用 30 次
- A State-Distribution Matching Approach to Non-Episodic Reinforcement LearningArchit Sharma, Rehaan Ahmad, Chelsea FinnICML 2022 · 被引用 23 次
- SOMBRL: Scalable and Optimistic Model-Based RLBhavya Sukhija, Lenart Treven, Carmelo Sferrazza, Florian Dörfler 等NeurIPS 2025 · 被引用 9 次
- NeoRL: Efficient Exploration for Nonepisodic RLBhavya Sukhija, Lenart Treven, Florian Dörfler, Stelian Coros 等NeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper11
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher 等ICML 2020 · 被引用 178 次
相关 Paper
- Demonstration-free Autonomous Reinforcement Learning via Implicit and Bidirectional CurriculumJigang Kim, Daesol Cho, H. Jin KimICML 2023 · 被引用 4 次
- Autonomous Reinforcement Learning via Subgoal CurriculaArchit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman 等NeurIPS 2021 · 被引用 41 次
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski 等NeurIPS 2021 · 被引用 152 次
- Robust and Scalable Autonomous Reinforcement Learning in Irreversible EnvironmentsSang-Hyun LeeNeurIPS 2025 · 被引用 1 次
- Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement LearningEvan Zheran Liu, Sahaana Suri, Tong Mu, Allan Zhou 等ICML 2023 · 被引用 4 次
