Performative Reinforcement Learning
Debmalya Mandal, Stelios Triantafyllou, Goran Radanovic
摘要
In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining approaches to a performatively stable policy. In the finite sample regime, these approaches repeatedly solve for a saddle point of a convex-concave objective, which estimates the Lagrangian of a regularized version of the reinforcement learning problem. In this paper, we aim to extend such repeated retraining approaches, enabling them to operate under corrupted data. More specifically, we consider Huber's ε-contamination model, where an ε fraction of data points is corrupted by arbitrary adversarial noise. We propose a repeated retraining approach based on convex-concave optimization under corrupted gradients and a novel problem-specific robust mean estimator for the gradients. We prove that our approach exhibits last-iterate convergence to an approximately stable policy, with the approximation error linear in √ε. We experimentally demonstrate the importance of accounting for corruption in performative reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Bilevel Optimization with Coupled Decision-Dependent DistributionsSongtao LuICML 2023 · 被引用 15 次
- Zero-Regret Performative Prediction Under Inequality ConstraintsWenjing Yan, Xuanyu CaoNeurIPS 2023 · 被引用 15 次
- Performative Control for Linear Dynamical SystemsSongfu Cai, Fei Han, Xuanyu CaoNeurIPS 2024 · 被引用 6 次
- Causal Inference out of Control: Estimating Performativity without Treatment RandomizationGary Cheng, Moritz Hardt, Celestine Mendler-DünnerICML 2024 · 被引用 4 次
- Decentralized Noncooperative Games with Coupled Decision-Dependent DistributionsWenjing Yan, Xuanyu CaoNeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper14
- Performative PredictionJuan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz HardtICML 2020 · 被引用 422 次
- Linear Last-iterate Convergence in Constrained Saddle-point OptimizationChen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, Haipeng LuoICLR 2021 · 被引用 146 次
- Dynamic Regret of Convex and Smooth FunctionsPeng Zhao, Yu-Jie Zhang, Lijun Zhang, Zhi-Hua ZhouNeurIPS 2020 · 被引用 136 次
- Outside the Echo Chamber: Optimizing the Performative RiskJohn Miller, Juan C. Perdomo, Tijana ZrnicICML 2021 · 被引用 128 次
- How to Learn when Data Reacts to Your Model: Performative Gradient DescentZachary Izzo, Lexing Ying, James ZouICML 2021 · 被引用 97 次
相关 Paper
- On Corruption-Robustness in Performative Reinforcement LearningVasilis Pollatos, Debmalya Mandal, Goran RadanovicAAAI 2025 · 被引用 6 次
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 被引用 43 次
- Outlier-Robust Optimal TransportDebarghya Mukherjee, Aritra Guha, Justin M. Solomon, Yuekai Sun 等ICML 2021 · 被引用 57 次
- Exact Policy Recovery in Offline RL with Both Heavy-Tailed Rewards and Data CorruptionYiding Chen, Xuezhou Zhang, Qiaomin Xie, Xiaojin ZhuAAAI 2024 · 被引用 2 次
- Adversarially Robust Change Point DetectionMengchu Li, Yi YuNeurIPS 2021 · 被引用 19 次
