Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
Erik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen, Scott Emmons, Stuart J. Russell
Abstract
Do neural networks learn to implement algorithms such as look-ahead or search"in the wild"? Or do they rely purely on collections of simple heuristics? We present evidence of learned look-ahead in the policy network of Leela Chess Zero, the currently strongest neural chess engine. We find that Leela internally represents future optimal moves and that these representations are crucial for its final output in certain board states. Concretely, we exploit the fact that Leela is a transformer that treats every chessboard square like a token in language models, and give three lines of evidence (1) activations on certain squares of future moves are unusually important causally; (2) we find attention heads that move important information"forward and backward in time,"e.g., from squares of future moves to squares of earlier ones; and (3) we train a simple probe that can predict the optimal move 2 turns ahead with 92% accuracy (in board states where Leela finds a single best line). These findings are an existence proof of learned look-ahead in neural networks and might be a step towards a better understanding of their capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2fd4563-ecc0-4d50-a197-faf3269d8514Cited by top-tier papers9
- Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language ModelsTianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen et al.EMNLP 2024 · 20 citations
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 11 citations
- Chessformer: A Unified Architecture for Chess ModelingDaniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang et al.ICLR 2026 · 8 citations
- Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNNMohammad Taufeeque, Aaron David Tucker, Adam Gleave, Adrià Garriga-AlonsoICLR 2026 · 5 citations
- Interpreting Emergent Planning in Model-Free Reinforcement LearningThomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso et al.ICLR 2025
Builds on12
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
Related papers
- Amortized Planning with Large-Scale Transformers: A Case Study on ChessAnian Ruoss, Grégoire Delétang, Sourabh Medapati, Jordi Grau-Moya et al.NeurIPS 2024 · 57 citations
- Chess as a Testbed for Language Model State TrackingShubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin GimpelAAAI 2022 · 77 citations
- Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess TransformersAnna Mészáros, Patrik Reizinger, Ferenc HuszárICML 2026
- Mastering Board Games by External and Internal Planning with Language ModelsJohn Schultz, Jakub Adámek, Matej Jusup, Marc Lanctot et al.ICML 2025
- A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled EnvironmentRaanan Yehezkel Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo et al.ICML 2025
