Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation
Mohit Prashant, Arvind Easwaran, Suman Das, Michael Yuhas
摘要
An issue concerning the use of deep reinforcement learning (RL) agents is whether they can be trusted to perform reliably when deployed, as training environments may not reflect real-life environments. Anticipating instances outside their training scope, learning-enabled systems are often equipped with out-of-distribution (OOD) detectors that alert when a trained system encounters a state it does not recognize or in which it exhibits uncertainty. There exists limited work conducted on the problem of OOD detection within RL, with prior studies being unable to achieve a consensus on the definition of OOD execution within the context of RL. By framing our problem using a Markov Decision Process, we assume there is a transition distribution mapping each state-action pair to another state with some probability. Based on this, we consider the following definition of OOD execution within RL: A transition is OOD if its probability during real-life deployment differs from the transition distribution encountered during training. As such, we utilize conditional variational autoencoders (CVAE) to approximate the transition dynamics of the training environment and implement a conformity-based detector using reconstruction loss that is able to guarantee OOD detection with a pre-determined confidence level. We evaluate our detector by adapting existing benchmarks and compare it with existing OOD detection models for RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Input Complexity and Out-of-distribution Detection with Likelihood-based Generative ModelsJoan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia 等ICLR 2020 · 被引用 307 次
- Likelihood Regret: An Out-of-Distribution Detection Score For Variational Auto-encoderZhisheng Xiao, Qing Yan, Yali AmitNeurIPS 2020 · 被引用 234 次
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind 等ICML 2021 · 被引用 223 次
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng 等ICLR 2022 · 被引用 173 次
相关 Paper
- Density-driven Regularization for Out-of-distribution DetectionWenjian Huang, Hao Wang, Jiahao Xia, Chengyan Wang 等NeurIPS 2022 · 被引用 17 次
- VOS: Learning What You Don't Know by Virtual Outlier SynthesisXuefeng Du, Zhaoning Wang, Mu Cai, Yixuan LiICLR 2022 · 被引用 417 次
- Learning to Shape In-distribution Feature Space for Out-of-distribution DetectionYonggang Zhang, Jie Lu, Bo Peng, Zhen Fang 等NeurIPS 2024 · 被引用 33 次
- iDECODe: In-Distribution Equivariance for Conformal Out-of-Distribution DetectionRamneet Kaur, Susmit Jha, Anirban Roy, Sangdon Park 等AAAI 2022 · 被引用 53 次
- Contrastive Out-of-Distribution Detection for Pretrained TransformersWenxuan Zhou, Fangyu Liu, Muhao ChenEMNLP 2021 · 被引用 63 次
