Detecting Rewards Deterioration in Episodic Reinforcement Learning
Ido Greenberg, Shie Mannor
摘要
In many RL applications, once training ends, it is vital to detect any deterioration in the agent performance as soon as possible. Furthermore, it often has to be done without modifying the policy and under minimal assumptions regarding the environment. In this paper, we address this problem by focusing directly on the rewards and testing for degradation. We consider an episodic framework, where the rewards within each episode are not independent, nor identically-distributed, nor Markov. We present this problem as a multivariate mean-shift detection problem with possibly partial observations. We define the mean-shift in a way corresponding to deterioration of a temporal signal (such as the rewards), and derive a test for this problem with optimal statistical power. Empirically, on deteriorated rewards in control problems (generated using various environment modifications), the test is demonstrated to be more powerful than standard tests - often by orders of magnitude. We also suggest a novel Bootstrap mechanism for False Alarm Rate control (BFAR), applicable to episodic (non-i.i.d) signal and allowing our test to run sequentially in an online manner. Our method does not rely on a learned model of the environment, is entirely external to the agent, and in fact can be applied to detect changes or drifts in any episodic signal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Train Hard, Fight Easy: Robust Meta Reinforcement LearningIdo Greenberg, Shie Mannor, Gal Chechik, Eli A. MeiromNeurIPS 2023 · 被引用 15 次
- e-COP : Episodic Constrained Optimization of PoliciesAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil SinglaNeurIPS 2024 · 被引用 2 次
- Novelty Detection in Reinforcement Learning with World ModelsGeigh Zollicoffer, Kenneth Eaton, Jonathan C. Balloch, Julia M. Kim 等ICML 2025
- Individualized Dosing Dynamics via Neural Eigen DecompositionStav Belogolovsky, Ido Greenberg, Danny Eytan, Shie MannorNeurIPS 2023
它引用的顶会 Paper5
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- Deployment-Efficient Reinforcement Learning via Model-Based Offline OptimizationTatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum 等ICLR 2021 · 被引用 166 次
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny 等ICLR 2020 · 被引用 99 次
- Bandits with Adversarial ScalingThodoris Lykouris, Vahab S. Mirrokni, Renato Paes LemeICML 2020 · 被引用 14 次
相关 Paper
- A Robust Test for the Stationarity Assumption in Sequential Decision MakingJitao Wang, Chengchun Shi, Zhenke WuICML 2023 · 被引用 9 次
- Anytime Detection of Strategic Deviations in Multi-Agent SystemsEtienne Gauthier, Francis Bach, Michael JordanICML 2026 · 被引用 2 次
- Testing For Distribution Shifts with Conditional Conformal Test MartingalesShalev Shaer, Yarin Bar, Drew Prinster, Yaniv RomanoICML 2026 · 被引用 1 次
- WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal MartingalesDrew Prinster, Xing Han, Anqi Liu, Suchi SariaICML 2025
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 被引用 36 次
