Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess Transformers
Anna Mészáros, Patrik Reizinger, Ferenc Huszár
Abstract
Modern decision transformers, trained similarly to LLMs, can achieve strong in-distribution performance in complex sequential domains like chess, but it remains unclear to what extent they reason systematically about rules and strategy. We study the reasoning capabilities of a 270M-parameter chess transformer trained via behavior cloning on standard chess. To investigate its abilities, we construct out-of-distribution test sets ---including board states and variants never seen during training---designed to reveal failures of systematic generalization. Our analysis shows that the model exhibits robust rule-based reasoning, consistently generating legal moves in novel configurations, but its strategic reasoning is more limited. The model generates high-quality moves on curated OOD puzzles and shows basic strategy adaptation in full games. It underperforms symbolic AI algorithms that rely on explicit search, although the performance gap is smaller when playing against human users on Lichess. Moreover, the training dynamics reveals distinct phases in how the model learns to respect the fundamental constraints, suggesting an emergent compositional understanding of the game.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton et al.NeurIPS 2025 · 507 citations
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin et al.ICLR 2024 · 189 citations
- Chess as a Testbed for Language Model State TrackingShubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin GimpelAAAI 2022 · 77 citations
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas et al.ICLR 2023 · 60 citations
Related papers
- Amortized Planning with Large-Scale Transformers: A Case Study on ChessAnian Ruoss, Grégoire Delétang, Sourabh Medapati, Jordi Grau-Moya et al.NeurIPS 2024 · 57 citations
- ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language ModelsJincheng Liu, Sijun He, Jingjing Wu, Xiangsen Wang et al.ACL 2026 · 1 citation
- Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann et al.ACL 2025 · 1 citation
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen et al.NeurIPS 2024 · 42 citations
- Can Transformers Reason Logically? A Study in SAT SolvingLeyan Pan, Vijay Ganesh, Jacob D. Abernethy, Chris Esposo et al.ICML 2025
