Lune

ICML2025顶会

Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement Learning

Austin A. Nguyen, Anri Gu, Michael P. Wellman

出版方
2025年份

摘要

Iterative extension of empirical game models through deep reinforcement learning (RL) has proved an effective approach for finding equilibria in complex games. When multiple equilibria exist, we may have preferences among solutions. We address this equilibrium selection issue in the context of Policy Space Response Oracles (PSRO), a flexible game-solving framework based on deep RL, by skewing strategy generation towards higher-welfare solutions. At each iteration, we create an exploration policy that imitates high welfare-yielding behavior and train a response to the current solution, regularized to be similar to the exploration policy. With no additional simulation expense, our approach, named Ex 2 PSRO, tends to find higher welfare equilibria than vanilla PSRO in two benchmarks: a sequential bargaining game and a social dilemma game. Further experiments demonstrate Ex 2 PSRO's composability with other PSRO variants and illuminate the relationship between exploration policy choice and algorithmic performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f1f4fe2a-e7ee-43ec-a391-9627308cca81

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖