Chess as a Testbed for Language Model State Tracking
Shubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin Gimpel
摘要
Transformer language models have made tremendous strides in natural language understanding tasks. However, the complexity of natural language makes it challenging to ascertain how accurately these models are tracking the world state underlying the text. Motivated by this issue, we consider the task of language modeling for the game of chess. Unlike natural language, chess notations describe a simple, constrained, and deterministic domain. Moreover, we observe that the appropriate choice of chess notation allows for directly probing the world state, without requiring any additional probing-related machinery. We find that: (a) With enough training data, transformer language models can learn to track pieces and predict legal moves with high accuracy when trained solely on move sequences. (b) For small training sets providing access to board state information during training can yield significant improvements. (c) The success of transformer language models is dependent on access to the entire game history i.e. “full attention”. Approximating this full attention results in a significant performance drop. We propose this testbed as a benchmark for future work on the development and analysis of transformer language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Evaluating the World Model Implicit in a Generative ModelKeyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg 等NeurIPS 2024 · 被引用 166 次
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 被引用 157 次
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Amortized Planning with Large-Scale Transformers: A Case Study on ChessAnian Ruoss, Grégoire Delétang, Sourabh Medapati, Jordi Grau-Moya 等NeurIPS 2024 · 被引用 57 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod 等ACL 2020 · 被引用 21 次
相关 Paper
- Verification of the Implicit World Model in a Generative Model via Adversarial SequencesAndrás Balogh, Márk JelasityICLR 2026 · 被引用 1 次
- Chessformer: A Unified Architecture for Chess ModelingDaniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang 等ICLR 2026 · 被引用 8 次
- Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess TransformersAnna Mészáros, Patrik Reizinger, Ferenc HuszárICML 2026
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen 等NeurIPS 2024 · 被引用 42 次
- A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled EnvironmentRaanan Yehezkel Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo 等ICML 2025
