Faster Game Solving via Hyperparameter Schedules
Naifeng Zhang, Stephen Marcus McAleer, Tuomas Sandholm
摘要
Counterfactual regret minimization (CFR) algorithms are a foundational class of methods for solving imperfect-information games, with the time average of their iterates converging to a Nash equilibrium in two-player zero-sum games. Prior state-of-the-art variants, Discounted CFR (DCFR) and Predictive CFR+ (PCFR+), achieved the fastest known practical performance by improving convergence rates over vanilla CFR through discounting early iterations with a fixed discounting scheme. More recently, Dynamic DCFR (DDCFR) introduced agent-learned dynamic discounting schemes to further accelerate convergence, at the cost of increased complexity. To address this, we propose Hyperparameter Schedules (HSs), a remarkably simple, training-free framework that dynamically adjusts CFR discounting over time. HSs aggressively downweight early updates and gradually transition to trusting late-stage strategies, leading to substantially faster convergence with only a few lines of code modifications. We show that HSs derived from just three small extensive-form games generalize effectively to 17 diverse games (including large-scale realistic poker) in both extensive-form and normal-form settings, without any game-specific tuning. Our method establishes a new state of the art for solving two-player zero-sum games.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Convergence of Regret Matching in Potential Games and Constrained OptimizationIoannis Anagnostides, Emanuel Tewolde, Brian Hu Zhang, Ioannis Panageas 等ICLR 2026 · 被引用 6 次
- Faster Game Solving via Asymmetry of Step SizesLinjian Meng, Tianpei Yang, Youzhi Zhang, Zhenxing Ge 等AAAI 2026
它引用的顶会 Paper4
- Faster Game Solving via Predictive Blackwell Approachability: Connecting Regret Matching and Mirror DescentGabriele Farina, Christian Kroer, Tuomas SandholmAAAI 2021 · 被引用 91 次
- Equilibrium Finding in Normal-Form Games via Greedy Regret MinimizationHugh Zhang, Adam Lerer, Noam BrownAAAI 2022 · 被引用 15 次
- Dynamic Discounted Counterfactual Regret MinimizationHang Xu, Kai Li, Haobo Fu, Qiang Fu 等ICLR 2024 · 被引用 7 次
- ESCHER: Eschewing Importance Sampling in Games by Computing a History Value Function to Estimate RegretStephen Marcus McAleer, Gabriele Farina, Marc Lanctot, Tuomas SandholmICLR 2023 · 被引用 1 次
相关 Paper
- AutoCFR: Learning to Design Counterfactual Regret Minimization AlgorithmsHang Xu, Kai Li, Haobo Fu, Qiang Fu 等AAAI 2022 · 被引用 12 次
- Posterior sampling for multi-agent reinforcement learning: solving extensive games with imperfect informationYichi Zhou, Jialian Li, Jun ZhuICLR 2020 · 被引用 18 次
- Preference-CFR: Beyond Nash Equilibrium for Better Game StrategiesQi Ju, Thomas Tellier, Meng Sun, Zhemei Fang 等ICML 2025
- Deep (Predictive) Discounted Counterfactual Regret MinimizationHang Xu, Kai Li, Haobo Fu, Qiang Fu 等AAAI 2026
- Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious PlayQi Ju, Falin Hei, Ting Feng, Dengbing Yi 等NeurIPS 2024 · 被引用 7 次
