Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
Mark Rofin, Jalal Naghiyev, Michael Hahn
Abstract
Trained Transformers have been shown to compute abstract features that appear redundant for predicting the immediate next token. We identify which components of the gradient signal from the next-token prediction objective give rise to this phenomenon, and we propose a method to estimate the influence of those components on the emergence of specific features. After validating our approach on toy tasks, we use it to interpret the origins of the world model in OthelloGPT and syntactic features in a small language model. Finally, we apply our framework to a pretrained LLM, showing that features with extremely high or low influence on future tokens tend to be related to formal reasoning domains such as code. Overall, our work takes a step toward understanding hidden features of Transformers through the lens of their development during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd950f91-0571-4de4-bb46-2b7e7b9f8b00Builds on18
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
Related papers
- Transformers learn factored representationsAdam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow et al.ICML 2026 · 2 citations
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 34 citations
- Pre-trained Large Language Models Use Fourier Features to Compute AdditionTianyi Zhou, Deqing Fu, Vatsal Sharan, Robin JiaNeurIPS 2024 · 48 citations
- Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsByung-Doh Oh, William SchulerACL 2023 · 1 citation
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
