AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement Learning
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li, Junliang Xing
Abstract
Heads-up no-limit Texas hold'em (HUNL) is the quintessential game with imperfect information. Representative prior works like DeepStack and Libratus heavily rely on counterfactual regret minimization (CFR) and its variants to tackle HUNL. However, the prohibitive computation cost of CFR iteration makes it difficult for subsequent researchers to learn the CFR model in HUNL and apply it in other practical applications. In this work, we present AlphaHoldem, a highperformance and lightweight HUNL AI obtained with an endto-end self-play reinforcement learning framework. The proposed framework adopts a pseudo-siamese architecture to directly learn from the input state information to the output actions by competing the learned model with its different historical versions. The main technical contributions include a novel state representation of card and betting information, a multi-task self-play training loss function, and a new model evaluation and selection metric to generate the final model. In a study involving 100,000 hands of poker, AlphaHoldem defeats Slumbot and DeepStack using only one PC with three days training. At the same time, AlphaHoldem only takes 2.9 milliseconds for each decision-making using only a single GPU, more than 1,000 times faster than DeepStack. We release the history data among among AlphaHoldem, Slumbot, and top human professionals in the author's GitHub repository to facilitate further studies in this direction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45c1ac10-fe81-4341-b304-9a3cb37ba957Cited by top-tier papers8
- EAVS: Edge-assisted Adaptive Video Streaming with Fine-grained Serverless PipelinesBiao Hou, Song Yang, Fernando A. Kuipers, Lei Jiao et al.INFOCOM 2023 · 26 citations
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter et al.ICLR 2024 · 19 citations
- How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool UseMinhua Lin, Enyan Dai, Hui Liu, Xianfeng Tang et al.ICLR 2026 · 9 citations
- Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?Renye Yan, Jikang Cheng, Shikun Sun, Yi Sun et al.CVPR 2026 · 9 citations
- RL-CFR: Improving Action Abstraction for Imperfect Information Extensive-Form Games with Reinforcement LearningBoning Li, Zhixuan Fang, Longbo HuangICML 2024 · 6 citations
Builds on3
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi et al.AAAI 2020 · 395 citations
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 205 citations
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini et al.ICLR 2021 · 52 citations
Related papers
- Efficient Online Pruning and Abstraction for Imperfect Information Extensive-Form GamesBoning Li, Longbo HuangICLR 2025
- Deep (Predictive) Discounted Counterfactual Regret MinimizationHang Xu, Kai Li, Haobo Fu, Qiang Fu et al.AAAI 2026
- No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained EmbeddingYanchang Fu, Shengda Liu, Pei Xu, Kaiqi HuangAAAI 2026
- An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form GamesLinjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An et al.AAAI 2023 · 8 citations
- Equilibrium Refinement for the Age of Machines: The One-Sided Quasi-Perfect EquilibriumGabriele Farina, Tuomas SandholmNeurIPS 2021 · 4 citations
