Faster Game Solving via Asymmetry of Step Sizes
Linjian Meng, Tianpei Yang, Youzhi Zhang, Zhenxing Ge, Yang Gao
摘要
Counterfactual Regret Minimization (CFR) algorithms are widely used to compute a Nash equilibrium (NE) in twoplayer zero-sum imperfect-information extensive-form games (IIGs). Among them, Predictive CFR + (PCFR + ) is particularly powerful, achieving an exceptionally fast empirical convergence rate via the prediction in many games. However, the empirical convergence rate of PCFR + would significantly degrade if the prediction is inaccurate, leading to unstable performance on certain IIGs. To enhance the robustness of PCFR + , we propose Asymmetric PCFR + (APCFR + ), which employs an adaptive asymmetry of step sizes between the updates of implicit and explicit accumulated counterfactual regrets to mitigate the impact of the prediction inaccuracy on convergence. We present a theoretical analysis demonstrating why APCFR + can enhance the robustness. To the best of our knowledge, we are the first to propose the asymmetry of step sizes, a simple yet novel technique that effectively improves the robustness of PCFR + . Then, to reduce the difficulty of implementing APCFR + caused by the adaptive asymmetry, we propose a simplified version of APCFR + called Simple APCFR + (SAPCFR + ), which uses a fixed asymmetry of step sizes to enable only a single-line modification compared to original PCFR + . Experimental results on five standard IIG benchmarks and two heads-up no-limit Texas Hold'em (HUNL) Subagems show that (i) both APCFR + and SAPCFR + outperform PCFR + in most of the tested games, (ii) SAPCFR + achieves a comparable empirical convergence rate with APCFR + , and (iii) our approach can be generalized to improve other CFR algorithms, e.g., Discount CFR (DCFR).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Linear Last-iterate Convergence in Constrained Saddle-point OptimizationChen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, Haipeng LuoICLR 2021 · 被引用 146 次
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei 等ICML 2021 · 被引用 102 次
- Faster Game Solving via Predictive Blackwell Approachability: Connecting Regret Matching and Mirror DescentGabriele Farina, Christian Kroer, Tuomas SandholmAAAI 2021 · 被引用 91 次
- Regret Matching+: (In)Stability and Fast Convergence in GamesGabriele Farina, Julien Grand-Clément, Christian Kroer, Chung-Wei Lee 等NeurIPS 2023 · 被引用 22 次
- AutoCFR: Learning to Design Counterfactual Regret Minimization AlgorithmsHang Xu, Kai Li, Haobo Fu, Qiang Fu 等AAAI 2022 · 被引用 12 次
相关 Paper
- Faster Game Solving via Hyperparameter SchedulesNaifeng Zhang, Stephen Marcus McAleer, Tuomas SandholmAAAI 2026 · 被引用 6 次
- Preference-CFR: Beyond Nash Equilibrium for Better Game StrategiesQi Ju, Thomas Tellier, Meng Sun, Zhemei Fang 等ICML 2025
- Dynamic Discounted Counterfactual Regret MinimizationHang Xu, Kai Li, Haobo Fu, Qiang Fu 等ICLR 2024 · 被引用 7 次
- Lazy-CFR: fast and near-optimal regret minimization for extensive games with imperfect informationYichi Zhou, Tongzheng Ren, Jialian Li, Dong Yan 等ICLR 2020 · 被引用 15 次
- An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form GamesLinjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An 等AAAI 2023 · 被引用 8 次
