Performative Policy Gradient: Optimality in Performative Reinforcement Learning
Debabrota Basu, Udvas Das, Brahim Driss, Uddalak Mukherjee
Abstract
Post-deployment machine learning algorithms often influence the environments that they act in, and thus, performatively shift the underlying dynamics that the standard Reinforcement Learning (RL) ignores. While designing optimal algorithms in this performative setting has been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and introduce the Performative Policy Gradient algorithm (PePG). PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9be229ea-b212-43db-943e-9b3b5e437062Builds on18
- Performative PredictionJuan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz HardtICML 2020 · 422 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Stochastic Optimization for Performative PredictionCelestine Mendler-Dünner, Juan C. Perdomo, Tijana Zrnic, Moritz HardtNeurIPS 2020 · 161 citations
- Outside the Echo Chamber: Optimizing the Performative RiskJohn Miller, Juan C. Perdomo, Tijana ZrnicICML 2021 · 128 citations
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 104 citations
Related papers
- How to Learn when Data Reacts to Your Model: Performative Gradient DescentZachary Izzo, Lexing Ying, James ZouICML 2021 · 97 citations
- Regret Minimization with Performative FeedbackMeena Jagadeesan, Tijana Zrnic, Celestine Mendler-DünnerICML 2022 · 41 citations
- Decentralized Noncooperative Games with Coupled Decision-Dependent DistributionsWenjing Yan, Xuanyu CaoNeurIPS 2024 · 4 citations
- Entropy-preserving reinforcement learningAleksei Petrenko, Ben Lipkin, Kevin Chen, Erik Wijmans et al.ICLR 2026 · 12 citations
- Stochastic Optimization Schemes for Performative Prediction with Nonconvex LossQiang Li, Hoi-To WaiNeurIPS 2024 · 18 citations
