Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation
Mohit Prashant, Arvind Easwaran, Suman Das, Michael Yuhas
Abstract
An issue concerning the use of deep reinforcement learning (RL) agents is whether they can be trusted to perform reliably when deployed, as training environments may not reflect real-life environments. Anticipating instances outside their training scope, learning-enabled systems are often equipped with out-of-distribution (OOD) detectors that alert when a trained system encounters a state it does not recognize or in which it exhibits uncertainty. There exists limited work conducted on the problem of OOD detection within RL, with prior studies being unable to achieve a consensus on the definition of OOD execution within the context of RL. By framing our problem using a Markov Decision Process, we assume there is a transition distribution mapping each state-action pair to another state with some probability. Based on this, we consider the following definition of OOD execution within RL: A transition is OOD if its probability during real-life deployment differs from the transition distribution encountered during training. As such, we utilize conditional variational autoencoders (CVAE) to approximate the transition dynamics of the training environment and implement a conformity-based detector using reconstruction loss that is able to guarantee OOD detection with a pre-determined confidence level. We evaluate our detector by adapting existing benchmarks and compare it with existing OOD detection models for RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e2ad2c8-b28d-47dc-aca3-d18388389c30Builds on12
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Input Complexity and Out-of-distribution Detection with Likelihood-based Generative ModelsJoan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia et al.ICLR 2020 · 307 citations
- Likelihood Regret: An Out-of-Distribution Detection Score For Variational Auto-encoderZhisheng Xiao, Qing Yan, Yali AmitNeurIPS 2020 · 234 citations
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind et al.ICML 2021 · 223 citations
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng et al.ICLR 2022 · 173 citations
Related papers
- Density-driven Regularization for Out-of-distribution DetectionWenjian Huang, Hao Wang, Jiahao Xia, Chengyan Wang et al.NeurIPS 2022 · 17 citations
- VOS: Learning What You Don't Know by Virtual Outlier SynthesisXuefeng Du, Zhaoning Wang, Mu Cai, Yixuan LiICLR 2022 · 417 citations
- Learning to Shape In-distribution Feature Space for Out-of-distribution DetectionYonggang Zhang, Jie Lu, Bo Peng, Zhen Fang et al.NeurIPS 2024 · 33 citations
- iDECODe: In-Distribution Equivariance for Conformal Out-of-Distribution DetectionRamneet Kaur, Susmit Jha, Anirban Roy, Sangdon Park et al.AAAI 2022 · 53 citations
- Contrastive Out-of-Distribution Detection for Pretrained TransformersWenxuan Zhou, Fangyu Liu, Muhao ChenEMNLP 2021 · 63 citations
