Solving Large-Scale Pursuit-Evasion Games Using Pre-trained Strategies
Shuxin Li, Xinrun Wang, Youzhi Zhang, Wanqi Xue, Jakub Cerný, Bo An
摘要
Pursuit-evasion games on graphs model the coordination of police forces chasing a fleeing felon in real-world urban settings, using the standard framework of imperfect-information extensive-form games (EFGs). In recent years, solving EFGs has been largely dominated by the Policy-Space Response Oracle (PSRO) methods due to their modularity, scalability, and favorable convergence properties. However, even these methods quickly reach their limits when facing large combinatorial strategy spaces of the pursuit-evasion games. To improve their efficiency, we integrate the pre-training and fine-tuning paradigm into the core module of PSRO -- the repeated computation of the best response. First, we pre-train the pursuer's policy base model against many different strategies of the evader. Then we proceed with the PSRO loop and fine-tune the pre-trained policy to attain the pursuer's best responses. The empirical evaluation shows that our approach significantly outperforms the baselines in terms of speed and scalability, and can solve even games on street maps of megalopolises with tens of thousands of crossroads -- a scale beyond the effective reach of previous methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Reevaluating Policy Gradient Methods for Imperfect-Information GamesMax Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen 等ICLR 2026 · 被引用 17 次
- Computing Optimal Nash Equilibria in Multiplayer GamesYouzhi Zhang, Bo An, Venkatramanan Siva SubrahmanianNeurIPS 2023 · 被引用 7 次
- DAG-Based Column Generation for Adversarial Team GamesYouzhi Zhang, Bo An, Daniel Dajun ZengICML 2024 · 被引用 6 次
- Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion GamesRunyu Lu, Peng Zhang, Ruochuan Shi, Yuanheng Zhu 等NeurIPS 2025 · 被引用 3 次
- R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial ObservabilityRunyu Lu, Ruochuan Shi, Yuanheng Zhu, Dongbin ZhaoICLR 2026
它引用的顶会 Paper3
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 被引用 98 次
- On the Effectiveness of Fine-tuning Versus Meta-reinforcement LearningMandi Zhao, Pieter Abbeel, Stephen JamesNeurIPS 2022 · 被引用 43 次
- NSGZero: Efficiently Learning Non-exploitable Policy in Large-Scale Network Security Games with Neural Monte Carlo Tree SearchWanqi Xue, Bo An, Chai Kiat YeoAAAI 2022 · 被引用 6 次
相关 Paper
- Iterative Empirical Game Solving via Single Policy Best ResponseMax Olan Smith, Thomas Anthony, Michael P. WellmanICLR 2021 · 被引用 23 次
- Tree-Based Stochastic Optimization for Solving Large-Scale Urban Network Security GamesShuxin Zhuang, Linjian Meng, Shuxin Li, Minming Li 等AAAI 2026
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
- Policy Space Diversity for Non-Transitive GamesJian Yao, Weiming Liu, Haobo Fu, Yaodong Yang 等NeurIPS 2023 · 被引用 28 次
- A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form GamesYudong Hu, Haoran Li, Congying Han, Tiande Guo 等ICML 2025
