Lune

ICLR2026顶会

Learning to Play Multi-Follower Bayesian Stackelberg Games

Gerson Personnat, Tao Lin, Safwan Hossain, David C. Parkes

2026年份
5被引次数
3顶会引用

摘要

In a multi-follower Bayesian Stackelberg game, a leader plays a mixed strategy over LL actions to which n≥1n\ge 1 followers, each having one of KK possible private types, best respond. The leader's optimal strategy depends on the distribution of the followers' private types. We study an online learning version of this problem: a leader interacts for TT rounds with nn followers with types sampled from an unknown distribution every round. The leader's goal is to minimize regret, defined as the difference between the cumulative utility of the optimal strategy and that of the actually chosen strategies. We design learning algorithms for the leader under different feedback settings. Under type feedback, where the leader observes the followers' types after each round, we design algorithms that achieve O(min⁡(Llog⁡(nKAT), nK)⋅T)O\big(\sqrt{\min(L\log(nKA T), ~ nK ) \cdot T} \big) regret for independent type distributions and O(min⁡(Llog⁡(nKAT), Kn)⋅T)O\big(\sqrt{\min(L\log(nKA T), ~ K^n ) \cdot T} \big) regret for general type distributions. Interestingly, those bounds do not grow with nn at a polynomial rate. Under action feedback, where the leader only observes the followers' actions, we design algorithms with O(min⁡(nLKLA2LLTlog⁡T, KnTlog⁡T))O( \min(\sqrt{ n^L K^L A^{2L} L T \log T}, ~ K^n\sqrt{ T } \log T ) ) regret. We also provide a lower bound of Ω(min⁡(L, nK)T)\Omega(\sqrt{\min(L, ~ nK)T}), almost matching the type-feedback upper bounds.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 55152b24-b3bc-4f6d-b0ef-924024a824b0

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖