Enhancing Chess Reinforcement Learning with Graph Representation
Tomas Rigaux, Hisashi Kashima
摘要
Mastering games is a hard task, as games can be extremely complex, and still fundamentally different in structure from one another. While the AlphaZero algorithm has demonstrated an impressive ability to learn the rules and strategy of a large variety of games, ranging from Go and Chess, to Atari games, its reliance on extensive computational resources and rigid Convolutional Neural Network (CNN) architecture limits its adaptability and scalability. A model trained to play on a Go board cannot be used to play on a smaller board, despite the similarity between the two Go variants. In this paper, we focus on Chess, and explore using a more generic Graph-based Representation of a game state, rather than a grid-based one, to introduce a more general architecture based on Graph Neural Networks (GNN). We also expand the classical Graph Attention Network (GAT) layer to incorporate edge-features, to naturally provide a generic policy output format. Our experiments, performed on smaller networks than the initial AlphaZero paper, show that this new architecture outperforms previous architectures with a similar number of parameters, being able to increase playing strength an order of magnitude faster. We also show that the model, when trained on a smaller variant of chess, is able to be quickly fine-tuned to play on regular chess, suggesting that this approach yields promising generalization abilities. Our code is available at https://github.com/akulen/AlphaGateau.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Deep Reinforcement Learning for General Game PlayingAdrian Goldwaser, Michael ThielscherAAAI 2020 · 被引用 46 次
- Scaling Laws for a Multi-Agent Reinforcement Learning ModelOren Neumann, Claudius GrosICLR 2023 · 被引用 3 次
- Are AlphaZero-like Agents Robust to Adversarial Perturbations?Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai 等NeurIPS 2022 · 被引用 15 次
- Evaluation beyond Task Performance: Analyzing Concepts in AlphaZero in HexCharles Lovering, Jessica Zosa Forde, George Konidaris, Ellie Pavlick 等NeurIPS 2022 · 被引用 13 次
- Multi-Agent Actor-Critic with Hierarchical Graph Attention NetworkHeechang Ryu, Hayong Shin, Jinkyoo ParkAAAI 2020 · 被引用 143 次
