Prioritizing Samples in Reinforcement Learning with Reducible Loss
Shivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi Kahou
摘要
Most reinforcement learning algorithms take advantage of an experience replay buffer to repeatedly train on samples the agent has observed in the past. Not all samples carry the same amount of significance and simply assigning equal importance to each of the samples is a naive strategy. In this paper, we propose a method to prioritize samples based on how much we can learn from a sample. We define the learn-ability of a sample as the steady decrease of the training loss associated with this sample over time. We develop an algorithm to prioritize samples with high learn-ability, while assigning lower priority to those that are hard-to-learn, typically caused by noise or stochasticity. We empirically show that across multiple domains our method is more robust than random sampling and also better than just prioritizing with respect to the training loss, i.e. the temporal difference loss, which is used in prioritized experience replay. The code to reproduce our experiments can be found here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu 等ICLR 2026 · 被引用 51 次
- Selective Learning for Deep Time Series ForecastingYisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li 等NeurIPS 2025 · 被引用 10 次
- REDUCR: Robust Data Downsampling using Class Priority ReweightingWilliam Bankes, George Hughes, Ilija Bogunovic, Zi WangNeurIPS 2024 · 被引用 5 次
- Off-policy Reinforcement Learning with Model-based Exploration AugmentationLikun Wang, Xiangteng Zhang, Yinuo Wang, Guojian Zhan 等NeurIPS 2025 · 被引用 3 次
- Looking Backward: Retrospective Backward Synthesis for Goal-Conditioned GFlowNetsHaoran He, Can Chang, Huazhe Xu, Ling PanICLR 2025
它引用的顶会 Paper11
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma 等ICML 2022 · 被引用 237 次
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 被引用 125 次
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 被引用 124 次
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 被引用 85 次
相关 Paper
- Regret Minimization Experience Replay in Off-Policy Reinforcement LearningXu-Hui Liu, Zhenghai Xue, Jing-Cheng Pang, Shengyi Jiang 等NeurIPS 2021 · 被引用 51 次
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 被引用 211 次
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 被引用 3 次
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 被引用 18 次
- Learning to Sample with Local and Global Contexts in Experience Replay BufferYoungmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang 等ICLR 2021 · 被引用 19 次
