Lune

ICML2025Top-tier venue

Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback

Canzhe Zhao, Yutian Cheng, Jing Dong, Baoxiang Wang, Shuai Li

2025Year

Abstract

We investigate learning approximate Nash equilibrium (NE) policy profiles in two-player zerosum imperfect information extensive-form games (IIEFGs) with last-iterate convergence guarantees. Existing algorithms either rely on full-information feedback or provide only asymptotic convergence rates. In contrast, we focus on the bandit feedback setting, where players receive feedback solely from the rewards associated with the experienced information set and action pairs in each episode. Our proposed algorithm employs a negentropy regularizer weighted by a "virtual transition" over the information set-action space to facilitate an efficient approximate policy update. Through a carefully designed virtual transition and leveraging the entropy regularization technique, we demonstrate finite-time last-iterate convergence to the NE with a rate of O(k -1 /8 ) under bandit feedback in each episode k. Empirical evaluations across various IIEFG instances show its competitive performance compared to baseline methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 78436ed8-56d2-4962-93fe-38814efd36af

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines