What Reward Structure Enables Efficient Sparse-Reward RL? A Proof-of-Concept with Policy-Aware Matrix Completion
Ibne Farabi Shihab, SANJEDA AKTER, Anuj Sharma
摘要
Sparse-reward reinforcement learning typically focuses on exploration, but we ask: can structural assumptions about reward functions themselves accelerate learning? We introduce Policy-Aware Matrix Completion (PAMC), which exploits low-rank structure in reward matrices while correcting for policy-induced sampling bias. PAMC combines three key components: a low-rank plus sparse reward model, inverse propensity weighting to handle Missing-Not-At-Random (MNAR) data, and confidence-gated abstention that falls back to intrinsic exploration when uncertain. We provide finite-sample theory showing that completion error scales as where ESS is the effective sample size under policy overlap . PAMC achieves strong empirical results at 10M steps (a sample-efficiency comparison): 4100250 return vs. 20050 for DrQ-v2 on Montezuma's Revenge, 78% vs. 65% success rate on MetaWorld-50, and 15% improvement over CQL on D4RL datasets. The method maintains 8% computational overhead while providing calibrated confidence intervals (95% empirical coverage). When structural assumptions are violated, PAMC gracefully degrades through increased abstention rather than catastrophic failure. Our approach demonstrates that reward structure exploitation can complement traditional exploration methods in sparse-reward domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
相关 Paper
- Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at RandomZiheng Wei, Annie Qu, Rui MiaoICML 2026
- Multi-User Reinforcement Learning with Low Rank RewardsDheeraj Mysore Nagaraj, Suhas S. Kowshik, Naman Agarwal, Praneeth Netrapalli 等ICML 2023 · 被引用 2 次
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning BenchmarksMuthukumar Pandaram, Jakob J. Hollenstein, David Drexel, Samuele Tosatto 等AAAI 2026 · 被引用 1 次
- Curiosity in Hindsight: Intrinsic Exploration in Stochastic EnvironmentsDaniel Jarrett, Corentin Tallec, Florent Altché, Thomas Mesnard 等ICML 2023 · 被引用 3 次
- Harnessing Structures for Value-Based Planning and Reinforcement LearningYuzhe Yang, Guo Zhang, Zhi Xu, Dina KatabiICLR 2020 · 被引用 39 次
