Partially Observable Reinforcement Learning with Memory Traces
Onno Eberhard, Michael Muehlebach, Claire Vernade
Abstract
Partially observable environments present a considerable computational challenge in reinforcement learning due to the need to consider long histories. Learning with a finite window of observations quickly becomes intractable as the window length grows. In this work, we introduce memory traces. Inspired by eligibility traces, these are compact representations of the history of observations in the form of exponential moving averages. We prove sample complexity bounds for the problem of offline on-policy evaluation that quantify the return errors achieved with memory traces for the class of Lipschitz continuous value estimates. We establish a close connection to the window approach, and demonstrate that, in certain environments, learning with memory traces is significantly more sample efficient. Finally, we underline the effectiveness of memory traces empirically in online reinforcement learning experiments for both value prediction and control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 746ec683-011f-498d-b502-a24e0576019aCited by top-tier papers5
- Kinaema: a recurrent sequence model for memory and pose in motionMert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci et al.NeurIPS 2025 · 3 citations
- Why Linear Recurrent Memory Works in Partially Observable Reinforcement LearningYike Zhao, Onno Eberhard, Malek khammassi, Ali Sayed et al.ICML 2026 · 2 citations
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei et al.ICML 2026
- Commit to the Bit: Reactive Reinforcement Learning Done RightOnno Eberhard, Claire Vernade, Michael MuehlebachICML 2026
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
Builds on4
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Provable Reinforcement Learning with a Short-Term MemoryYonathan Efroni, Chi Jin, Akshay Krishnamurthy, Sobhan MiryoosefiICML 2022 · 45 citations
- Mitigating Partial Observability in Sequential Decision Processes via the Lambda DiscrepancyCameron Allen, Aaron Kirtland, Ruo Yu Tao, Sam Lobel et al.NeurIPS 2024 · 12 citations
- Planning and Learning in Partially Observable Systems via Filter StabilityNoah Golowich, Ankur Moitra, Dhruv RohatgiSTOC 2023 · 4 citations
Related papers
- On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent SamplesMustafa O. Karabag, Ufuk TopcuAAAI 2023 · 6 citations
- Reusing Trajectories in Policy Gradients Enables Fast ConvergenceAlessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini et al.ICML 2026
- Represent to Control Partially Observed Systems: Representation Learning with Provable Sample EfficiencyLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2023
- Provably Efficient Reinforcement Learning in Partially Observable Dynamical SystemsMasatoshi Uehara, Ayush Sekhari, Jason D. Lee, Nathan Kallus et al.NeurIPS 2022 · 48 citations
- Real-Time Recurrent Learning using Trace Units in Reinforcement LearningEsraa Elelimy, Adam White, Michael Bowling, Martha WhiteNeurIPS 2024 · 15 citations
