Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration
Desik Rengarajan, Gargi Vaidya, Akshay Sarvesh, Dileep M. Kalathil, Srinivas Shakkottai
摘要
A major challenge in real-world reinforcement learning (RL) is the sparsity of reward feedback. Often, what is available is an intuitive but sparse reward function that only indicates whether the task is completed partially or fully. However, the lack of carefully designed, fine grain feedback implies that most existing RL algorithms fail to learn an acceptable policy in a reasonable time frame. This is because of the large number of exploration actions that the policy has to perform before it gets any useful feedback that it can learn from. In this work, we address this challenging problem by developing an algorithm that exploits the offline demonstration data generated by a sub-optimal behavior policy for faster and efficient online RL in such sparse reward settings. The proposed algorithm, which we call the Learning Online with Guidance Offline (LOGO) algorithm, merges a policy improvement step with an additional policy guidance step by using the offline demonstration data. The key idea is that by obtaining guidance from - not imitating - the offline data, LOGO orients its policy in the manner of the sub-optimal policy, while yet being able to learn beyond and approach optimality. We provide a theoretical analysis of our algorithm, and provide a lower bound on the performance improvement in each learning episode. We also extend our algorithm to the even more challenging incomplete observation setting, where the demonstration data contains only a censored version of the true state observation. We demonstrate the superior performance of our algorithm over state-of-the-art approaches on a number of benchmark environments with sparse rewards and censored state. Further, we demonstrate the value of our approach via implementing LOGO on a mobile robot for trajectory tracking and obstacle avoidance, where it shows excellent performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningMingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang 等EMNLP 2022 · 被引用 141 次
- ReDit: Reward Dithering for Improved LLM Policy OptimizationChenxing Wei, Jiarui Yu, Ying He, Hande Dong 等NeurIPS 2025 · 被引用 14 次
- Hybrid Policy Optimization from Imperfect DemonstrationsHanlin Yang, Chao Yu, Peng Sun, Siji ChenNeurIPS 2023 · 被引用 14 次
- Towards VM Rescheduling Optimization Through Deep Reinforcement LearningXianzhong Ding, Yunkai Zhang, Binbin Chen, Donghao Ying 等EuroSys 2025 · 被引用 10 次
- Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language ModelsYanjiang Liu, Shuheng Zhou, Yaojie Lu, Huijia Zhu 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper3
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki 等ICLR 2020 · 被引用 299 次
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil 等ICLR 2020 · 被引用 97 次
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 被引用 32 次
相关 Paper
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil 等NeurIPS 2022 · 被引用 2 次
- Accelerating Exploration with Unlabeled Prior DataQiyang Li, Jason Zhang, Dibya Ghosh, Amy Zhang 等NeurIPS 2023 · 被引用 21 次
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng 等AAAI 2025
- GUIDE: Real-Time Human-Shaped AgentsLingyu Zhang, Zhengran Ji, Nicholas R. Waytowich, Boyuan ChenNeurIPS 2024 · 被引用 9 次
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 被引用 84 次
