Influencing Long-Term Behavior in Multiagent Reinforcement Learning
Dong-Ki Kim, Matthew Riemer, Miao Liu, Jakob N. Foerster, Michael Everett, Chuangchuang Sun, Gerald Tesauro, Jonathan P. How
Abstract
The main challenge of multiagent reinforcement learning is the difficulty of learning useful policies in the presence of other simultaneously learning agents whose changing behaviors jointly affect the environment's transition and reward dynamics. An effective approach that has recently emerged for addressing this non-stationarity is for each agent to anticipate the learning of other agents and influence the evolution of future policies towards desirable behavior for its own benefit. Unfortunately, previous approaches for achieving this suffer from myopic evaluation, considering only a finite number of policy updates. As such, these methods can only influence transient future policies rather than achieving the promise of scalable equilibrium selection approaches that influence the behavior at convergence. In this paper, we propose a principled framework for considering the limiting policies of other agents as time approaches infinity. Specifically, we develop a new optimization objective that maximizes each agent's average reward by directly accounting for the impact of its behavior on the limiting set of policies that other agents will converge to. Our paper characterizes desirable solution concepts within this problem setting and provides practical approaches for optimizing over possible outcomes. As a result of our farsighted objective, we demonstrate better long-term performance than state-of-the-art baselines across a suite of diverse multiagent benchmark domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77f57c44-d1c5-48ae-bbb4-09bdb5d0a10cCited by top-tier papers8
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement LearningBoyin Liu, Zhiqiang Pu, Yi Pan, Jianqiang Yi et al.ICML 2023 · 34 citations
- Learning to Influence Human Behavior with Offline Reinforcement LearningJoey Hong, Sergey Levine, Anca D. DraganNeurIPS 2023 · 30 citations
- Continual Learning In Environments With Polynomial Mixing TimesMatthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj et al.NeurIPS 2022 · 18 citations
- On the Interplay between Social Welfare and Tractability of EquilibriaIoannis Anagnostides, Tuomas SandholmNeurIPS 2023 · 3 citations
Builds on5
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Learning to Incentivize Other Learning AgentsJiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag et al.NeurIPS 2020 · 105 citations
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement LearningDong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun et al.ICML 2021 · 66 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Continuous-Time Meta-Learning with Forward Mode DifferentiationTristan Deleu, David Kanaa, Leo Feng, Giancarlo Kerg et al.ICLR 2022 · 22 citations
Related papers
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar et al.NeurIPS 2021 · 105 citations
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
- Paths to Equilibrium in GamesBora Yongacoglu, Gürdal Arslan, Lacra Pavel, Serdar YükselNeurIPS 2024 · 2 citations
- Opponent Modeling based on Subgoal InferenceXiaopeng Yu, Jiechuan Jiang, Zongqing LuNeurIPS 2024 · 7 citations
- Maximizing utility in multi-agent environments by anticipating the behavior of other learnersAngelos Assos, Yuval Dagan, Constantinos DaskalakisNeurIPS 2024 · 16 citations
