When to Ask for Help: Proactive Interventions in Autonomous Reinforcement Learning
Annie Xie, Fahim Tajwar, Archit Sharma, Chelsea Finn
摘要
A long-term goal of reinforcement learning is to design agents that can autonomously interact and learn in the world. A critical challenge to such autonomy is the presence of irreversible states which require external assistance to recover from, such as when a robot arm has pushed an object off of a table. While standard agents require constant monitoring to decide when to intervene, we aim to design proactive agents that can request human intervention only when needed. To this end, we propose an algorithm that efficiently learns to detect and avoid states that are irreversible, and proactively asks for help in case the agent does enter them. On a suite of continuous control environments with unknown irreversible states, we find that our algorithm exhibits better sample-and intervention-efficiency compared to existing methods. Our code is publicly available at https://sites.google.com/view/proactive-interventions . Figure 1: Autonomous agents struggle to make progress without external interventions when they are stuck in an irreversible state. Reinforcement learning agents therefore need active monitoring throughout training to detect and intervene when the agent reaches an irreversible state. Enabling the agents to detect irreversible states and proactively request for help can substantially reduce the human monitoring required for training agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataFahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov 等ICML 2024 · 被引用 189 次
- SAFE: Multitask Failure Detection for Vision-Language-Action ModelsQiao Gu, Yuanliang Ju, Shengxiang Sun, Igor Gilitschenski 等NeurIPS 2025 · 被引用 103 次
- Tools Fail: Detecting Silent Errors in Faulty ToolsJimin Sun, So Yeon Min, Yingshan Chang, Yonatan BiskEMNLP 2024 · 被引用 4 次
- Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action ModelsShelly Francis-Meretzki, Mirco Mutti, Yaniv Romano, Aviv TamarICML 2026 · 被引用 2 次
- Robust and Scalable Autonomous Reinforcement Learning in Irreversible EnvironmentsSang-Hyun LeeNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper15
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
相关 Paper
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 等NeurIPS 2020 · 被引用 109 次
- There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement LearningNathan Grinsztajn, Johan Ferret, Olivier Pietquin, Philippe Preux 等NeurIPS 2021 · 被引用 23 次
- Autonomous Reinforcement Learning via Subgoal CurriculaArchit Sharma, Abhishek Gupta, Sergey Levine, Karol Hausman 等NeurIPS 2021 · 被引用 41 次
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention MechanismHaoyuan Cai, Zhenghao Peng, Bolei ZhouICML 2025
- Predictive Preference Learning from Human InterventionsHaoyuan Cai, Zhenghao Mark Peng, Bolei ZhouNeurIPS 2025 · 被引用 6 次
