Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning Algorithms
Liyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov, Lillian J. Ratliff
摘要
The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction as a two-player general-sum game with a leader-follower structure known as a Stackelberg game. Given this abstraction, we propose a meta-framework for Stackelberg actor-critic algorithms where the leader player follows the total derivative of its objective instead of the usual individual gradient. From a theoretical standpoint, we develop a policy gradient theorem for the refined update and provide a local convergence guarantee for the Stackelberg actor-critic algorithms to a local Stackelberg equilibrium. From an empirical standpoint, we demonstrate via simple examples that the learning dynamics we study mitigate cycling and accelerate convergence compared to the usual gradient dynamics given cost structures induced by actor-critic formulations. Finally, extensive experiments on OpenAI gym environments show that Stackelberg actor-critic algorithms always perform at least as well and often significantly outperform the standard actor-critic algorithm counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 被引用 176 次
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 被引用 156 次
- Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticTianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan 等ICML 2024 · 被引用 23 次
- Score Models for Offline Goal-Conditioned Reinforcement LearningHarshit Sikchi, Rohan Chitnis, Ahmed Touati, Alborz Geramifard 等ICLR 2024 · 被引用 16 次
- Belief-Enriched Pessimistic Q-Learning against Adversarial State PerturbationsXiaolin Sun, Zizhan ZhengICLR 2024 · 被引用 4 次
它引用的顶会 Paper7
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 被引用 144 次
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 被引用 137 次
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li 等AAAI 2020 · 被引用 113 次
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny 等ICLR 2020 · 被引用 99 次
- Characterizing the Gap Between Actor-Critic and Policy GradientJunfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale SchuurmansICML 2021 · 被引用 18 次
相关 Paper
- Learning in Stackelberg Mean Field Games: A Non-Asymptotic AnalysisSihan Zeng, Benjamin Patrick Evans, Sujay Bhatt, Leo Ardon 等NeurIPS 2025 · 被引用 1 次
- Oracles & Followers: Stackelberg Equilibria in Deep Multi-Agent Reinforcement LearningMatthias Gerstgrasser, David C. ParkesICML 2023 · 被引用 27 次
- Stackelberg Coupling of Online Representation Learning and Reinforcement LearningFernando Martinez, Tao Li, Yingdong Lu, Juntao ChenICLR 2026 · 被引用 1 次
- Rethinking Neural Network Learning Rates: A Stackelberg PerspectiveSihan Zeng, Sujay Bhatt, Sumitra GaneshICML 2026
- Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg GamesKai Wang, Lily Xu, Andrew Perrault, Michael K. Reiter 等AAAI 2022 · 被引用 29 次
