On Corruption-Robustness in Performative Reinforcement Learning
Vasilis Pollatos, Debmalya Mandal, Goran Radanovic
Abstract
In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining approaches to a performatively stable policy. In the finite sample regime, these approaches repeatedly solve for a saddle point of a convex-concave objective, which estimates the Lagrangian of a regularized version of the reinforcement learning problem. In this paper, we aim to extend such repeated retraining approaches, enabling them to operate under corrupted data. More specifically, we consider Huber's ε-contamination model, where an ε fraction of data points is corrupted by arbitrary adversarial noise. We propose a repeated retraining approach based on convex-concave optimization under corrupted gradients and a novel problem-specific robust mean estimator for the gradients. We prove that our approach exhibits last-iterate convergence to an approximately stable policy, with the approximation error linear in √ε. We experimentally demonstrate the importance of accounting for corruption in performative reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56d5d4ee-06c1-4cff-857c-4b049e9fa6ccCited by top-tier papers2
- Scalable Valuation of Human Feedback through Provably Robust Model AlignmentMasahiro Fujisawa, Masaki Adachi, Michael A. OsborneNeurIPS 2025 · 6 citations
- Performative Policy Gradient: Optimality in Performative Reinforcement LearningDebabrota Basu, Udvas Das, Brahim Driss, Uddalak MukherjeeICML 2026 · 2 citations
Builds on14
- Performative PredictionJuan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz HardtICML 2020 · 422 citations
- Linear Last-iterate Convergence in Constrained Saddle-point OptimizationChen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, Haipeng LuoICLR 2021 · 146 citations
- Dynamic Regret of Convex and Smooth FunctionsPeng Zhao, Yu-Jie Zhang, Lijun Zhang, Zhi-Hua ZhouNeurIPS 2020 · 136 citations
- Outside the Echo Chamber: Optimizing the Performative RiskJohn Miller, Juan C. Perdomo, Tijana ZrnicICML 2021 · 128 citations
- How to Learn when Data Reacts to Your Model: Performative Gradient DescentZachary Izzo, Lexing Ying, James ZouICML 2021 · 97 citations
Related papers
- Performative Reinforcement LearningDebmalya Mandal, Stelios Triantafyllou, Goran RadanovicICML 2023 · 6 citations
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 43 citations
- Outlier-Robust Optimal TransportDebarghya Mukherjee, Aritra Guha, Justin M. Solomon, Yuekai Sun et al.ICML 2021 · 57 citations
- Exact Policy Recovery in Offline RL with Both Heavy-Tailed Rewards and Data CorruptionYiding Chen, Xuezhou Zhang, Qiaomin Xie, Xiaojin ZhuAAAI 2024 · 2 citations
- Adversarially Robust Change Point DetectionMengchu Li, Yi YuNeurIPS 2021 · 19 citations
