Learning to Sample with Local and Global Contexts in Experience Replay Buffer
Youngmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang, Sung Ju Hwang
摘要
Experience replay, which enables the agents to remember and reuse experience from the past, has played a significant role in the success of off-policy reinforcement learning (RL). To utilize the experience replay efficiently, the existing sampling methods allow selecting out more meaningful experiences by imposing priorities on them based on certain metrics (e.g. TD-error). However, they may result in sampling highly biased, redundant transitions since they compute the sampling rate for each transition independently, without consideration of its importance in relation to other transitions. In this paper, we aim to address the issue by proposing a new learning-based sampling method that can compute the relative importance of transition. To this end, we design a novel permutation-equivariant neural architecture that takes contexts from not only features of each transition (local) but also those of others (global) as inputs. We validate our framework, which we refer to as Neural Experience Replay Sampler (NERS) 1 , on multiple benchmark tasks for both continuous and discrete control tasks and show that it can significantly improve the performance of various off-policy RL methods. Further analysis confirms that the improvements of the sample efficiency indeed are due to sampling diverse and meaningful transitions by NERS that considers both local and global contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Model-augmented Prioritized Experience ReplayYoungmin Oh, Jinwoo Shin, Eunho Yang, Sung Ju HwangICLR 2022 · 被引用 21 次
- Curious Replay for Model-based AdaptationIsaac Kauvar, Chris Doyle, Linqi Zhou, Nick HaberICML 2023 · 被引用 18 次
- Learning from Good Trajectories in Offline Multi-Agent Reinforcement LearningQi Tian, Kun Kuang, Furui Liu, Baoxiang WangAAAI 2023 · 被引用 14 次
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 被引用 3 次
- Replay Memory as An Empirical MDP: Combining Conservative Estimation with Experience ReplayHongming Zhang, Chenjun Xiao, Han Wang, Jun Jin 等ICLR 2023
它引用的顶会 Paper2
相关 Paper
- Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement LearningBrett Daley, Martha White, Christopher Amato, Marlos C. MachadoICML 2023 · 被引用 4 次
- Attentive Experience ReplayPeiquan Sun, Wengang Zhou, Houqiang LiAAAI 2020 · 被引用 62 次
- Prioritizing Samples in Reinforcement Learning with Reducible LossShivakanth Sujit, Somjit Nath, Pedro H. M. Braga, Samira Ebrahimi KahouNeurIPS 2023 · 被引用 36 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
- Predicting the Susceptibility of Examples to Catastrophic ForgettingGuy Hacohen, Tinne TuytelaarsICML 2025
