Understanding and Diagnosing Deep Reinforcement Learning
Ezgi Korkmaz
Abstract
Deep neural policies have recently been installed in a diverse range of settings, from biotechnology to automated financial systems. However, the utilization of deep neural networks to approximate the value function leads to concerns on the decision boundary stability, in particular, with regard to the sensitivity of policy decision making to indiscernible, non-robust features due to highly non-convex and complex deep neural manifolds. These concerns constitute an obstruction to understanding the reasoning made by deep neural policies, and their foundational limitations. Hence, it is crucial to develop techniques that aim to understand the sensitivities in the learnt representations of neural network policies. To achieve this we introduce a theoretically founded method that provides a systematic analysis of the unstable directions in the deep neural policy decision boundary across both time and space. Through experiments in the Arcade Learning Environment (ALE), we demonstrate the effectiveness of our technique for identifying correlated directions of instability, and for measuring how sample shifts remold the set of sensitive directions in the neural policy landscape. Most importantly, we demonstrate that state-of-the-art robust training techniques yield learning of disjoint unstable directions, with dramatically larger oscillations over time, when compared to standard training. We believe our results reveal the fundamental properties of the decision process made by reinforcement learning policies, and can help in constructing reliable and robust deep neural policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- UR² : Unify RAG and Reasoning through Reinforcement LearningWeitao Li, Boran Xiang, Xiaolong Wang, Jingyi Ren et al.ACL 2026 · 1 citation
- Counteractive RL: Rethinking Core Principles for Efficient and Scalable Deep Reinforcement LearningEzgi KorkmazNeurIPS 2025 · 1 citation
- How to Lose Inherent Counterfactuality in Reinforcement LearningEzgi KorkmazICLR 2026
- Principled Analysis of Deep Reinforcement Learning Evaluation and Design ParadigmsEzgi KorkmazAAAI 2026
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot et al.ICML 2020 · 103 citations
- Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPsEzgi KorkmazAAAI 2022 · 33 citations
Related papers
- Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust DecisionsEzgi Korkmaz, Jonah Brown-CohenICML 2023 · 16 citations
- Adversarial Robust Deep Reinforcement Learning Requires Redefining RobustnessEzgi KorkmazAAAI 2023 · 38 citations
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon et al.NeurIPS 2023 · 11 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- Finding separatrices of dynamical flows with Deep Koopman EigenfunctionsKabir V. Dabholkar, Omri BarakNeurIPS 2025 · 3 citations
