Detecting Rewards Deterioration in Episodic Reinforcement Learning
Ido Greenberg, Shie Mannor
Abstract
In many RL applications, once training ends, it is vital to detect any deterioration in the agent performance as soon as possible. Furthermore, it often has to be done without modifying the policy and under minimal assumptions regarding the environment. In this paper, we address this problem by focusing directly on the rewards and testing for degradation. We consider an episodic framework, where the rewards within each episode are not independent, nor identically-distributed, nor Markov. We present this problem as a multivariate mean-shift detection problem with possibly partial observations. We define the mean-shift in a way corresponding to deterioration of a temporal signal (such as the rewards), and derive a test for this problem with optimal statistical power. Empirically, on deteriorated rewards in control problems (generated using various environment modifications), the test is demonstrated to be more powerful than standard tests - often by orders of magnitude. We also suggest a novel Bootstrap mechanism for False Alarm Rate control (BFAR), applicable to episodic (non-i.i.d) signal and allowing our test to run sequentially in an online manner. Our method does not rely on a learned model of the environment, is entirely external to the agent, and in fact can be applied to detect changes or drifts in any episodic signal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84916f3c-1600-495b-8ac0-d31c261b6dbaCited by top-tier papers4
- Train Hard, Fight Easy: Robust Meta Reinforcement LearningIdo Greenberg, Shie Mannor, Gal Chechik, Eli A. MeiromNeurIPS 2023 · 15 citations
- e-COP : Episodic Constrained Optimization of PoliciesAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil SinglaNeurIPS 2024 · 2 citations
- Novelty Detection in Reinforcement Learning with World ModelsGeigh Zollicoffer, Kenneth Eaton, Jonathan C. Balloch, Julia M. Kim et al.ICML 2025
- Individualized Dosing Dynamics via Neural Eigen DecompositionStav Belogolovsky, Ido Greenberg, Danny Eytan, Shie MannorNeurIPS 2023
Builds on5
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Deployment-Efficient Reinforcement Learning via Model-Based Offline OptimizationTatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum et al.ICLR 2021 · 166 citations
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny et al.ICLR 2020 · 99 citations
- Bandits with Adversarial ScalingThodoris Lykouris, Vahab S. Mirrokni, Renato Paes LemeICML 2020 · 14 citations
Related papers
- A Robust Test for the Stationarity Assumption in Sequential Decision MakingJitao Wang, Chengchun Shi, Zhenke WuICML 2023 · 9 citations
- Anytime Detection of Strategic Deviations in Multi-Agent SystemsEtienne Gauthier, Francis Bach, Michael JordanICML 2026 · 2 citations
- Testing For Distribution Shifts with Conditional Conformal Test MartingalesShalev Shaer, Yarin Bar, Drew Prinster, Yaniv RomanoICML 2026 · 1 citation
- WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal MartingalesDrew Prinster, Xing Han, Anqi Liu, Suchi SariaICML 2025
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 36 citations
