The Geometry of Robust Value Functions
Kaixin Wang, Navdeep Kumar, Kuangqi Zhou, Bryan Hooi, Jiashi Feng, Shie Mannor
摘要
The space of value functions is a fundamental concept in reinforcement learning. Characterizing its geometric properties may provide insights for optimization and representation. Existing works mainly focus on the value space for Markov Decision Processes (MDPs). In this paper, we study the geometry of the robust value space for the more general Robust MDPs (RMDPs) setting, where transition uncertainties are considered. Specifically, since we find it hard to directly adapt prior approaches to RMDPs, we start with revisiting the non-robust case, and introduce a new perspective that enables us to characterize both the non-robust and robust value space in a similar fashion. The key of this perspective is to decompose the value space, in a state-wise manner, into unions of hypersurfaces. Through our analysis, we show that the robust value space is determined by a set of conic hypersurfaces, each of which contains the robust values of all policies that agree on one state. Furthermore, we find that taking only extreme points in the uncertainty set is sufficient to determine the robust value space. Finally, we discuss some other aspects about the robust value space, including its non-convexity and policy agreement on multiple states.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Policy Gradient in Robust MDPs with Global Convergence GuaranteeQiuhao Wang, Chin Pang Ho, Marek PetrikICML 2023 · 被引用 43 次
- Fast Bellman Updates for Wasserstein Distributionally Robust MDPsZhuodong Yu, Ling Dai, Shaohang Xu, Siyang Gao 等NeurIPS 2023 · 被引用 15 次
- Non-rectangular Robust MDPs with Normed Uncertainty SetsNavdeep Kumar, Adarsh Gupta, Maxence Mohamed Elfatihi, Giorgia Ramponi 等NeurIPS 2025 · 被引用 3 次
- Geometric Policy Iteration for Markov Decision ProcessesYue Wu, Jesús A. De LoeraKDD 2022 · 被引用 1 次
- The Value Function Semi-Algebraic Set in Partially Observable Markov Decision ProcessesRyan Anderson, Guido MontufarICML 2026
它引用的顶会 Paper1
相关 Paper
- Solving Robust Markov Decision Processes: Generic, Reliable, EfficientTobias Meggendorfer, Maximilian Weininger, Patrick WienhöftAAAI 2025
- Model-Free Robust Average-Reward Reinforcement LearningYue Wang, Alvaro Velasquez, George K. Atia, Ashley Prater-Bennette 等ICML 2023 · 被引用 25 次
- Efficient Solution and Learning of Robust Factored MDPsYannik Schnitzer, Alessandro Abate, David ParkerAAAI 2026 · 被引用 1 次
- Best-Effort Policies for Robust Markov Decision ProcessesAlessandro Abate, Thom Badings, Giuseppe De Giacomo, Francesco FabianoAAAI 2026
- Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPsEzgi KorkmazAAAI 2022 · 被引用 33 次
