Learning a Game by Paying the Agents
Brian Hu Zhang, Tao Lin, Yiling Chen, Tuomas Sandholm
Abstract
We study the problem of learning the utility functions of no-regret learning agents in a repeated normal-form game. Differing from most prior literature, we introduce a principal with the power to observe the agents playing the game, send agents signals, and give agents payments as a function of their actions. We show that the principal can, using a number of rounds polynomial in the size of the game, learn the utility functions of all agents to any desired precision ε > 0, for any no-regret learning algorithms of the agents. Our main technique is to formulate a zero-sum game between the principal and the agents, where the principal chooses strategies among the set of all payment functions to minimize the agent's payoff. Finally, we discuss implications for the problem of steering agents. We introduce, using our utility-learning algorithm as a subroutine, the first algorithm for steering arbitrary no-regret learning agents to a desired equilibrium without prior knowledge of their utility functions. * Equal contribution. Order determined randomly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f085ef1-2add-438a-9227-bbd9370d064dCited by top-tier papers2
- Inferring Heterogeneous Private Valuations from Offline Market Data via Entropic Risk-Sensitive Utility MaximizationXingyu Qian, Haoran YuAAAI 2026
- Optimal Robust Subsidy Policies for Irrational Agent in Principal-Agent MDPsBowen Hu, Yixin TaoICLR 2026
Builds on6
- Fast Last-Iterate Convergence of Learning in Games Requires Forgetful AlgorithmsYang Cai, Gabriele Farina, Julien Grand-Clément, Christian Kroer et al.NeurIPS 2024 · 24 citations
- Online Bayesian Persuasion Without a ClueFrancesco Bacchiocchi, Matteo Bollini, Matteo Castiglioni, Alberto Marchesi et al.NeurIPS 2024 · 9 citations
- Learning to Mitigate Externalities: the Coase Theorem with Hindsight RationalityAntoine Scheid, Aymeric Capitaine, Etienne Boursier, Eric Moulines et al.NeurIPS 2024 · 7 citations
- Learning to Steer Learners in GamesYizhou Zhang, Yian Ma, Eric MazumdarICML 2025
- Single-agent Poisoning Attacks Suffice to Ruin Multi-Agent LearningFan Yao, Yuwei Cheng, Ermin Wei, Haifeng XuICLR 2025
Related papers
- Generalized Principal-Agent Problem with a Learning AgentTao Lin, Yiling ChenICLR 2025
- Learning Optimal Contracts: How to Exploit Small Action SpacesFrancesco Bacchiocchi, Matteo Castiglioni, Alberto Marchesi, Nicola GattiICLR 2024 · 21 citations
- Incentivized Learning in Principal-Agent Bandit GamesAntoine Scheid, Daniil Tiapkin, Etienne Boursier, Aymeric Capitaine et al.ICML 2024 · 17 citations
- Contracting with a Learning AgentGuru Guruganesh, Yoav Kolumbus, Jon Schneider, Inbal Talgam-Cohen et al.NeurIPS 2024 · 38 citations
- Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent ArrivalsJunyan Liu, Arnab Maiti, Artin Tajdini, Kevin Jamieson et al.ICML 2025
