Hidden Breakthroughs in Language Model Training
Sara Kangaslahti, Elan Rosenfeld, Naomi Saphra
Abstract
Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. These breakthroughs enable a deeper understanding of the model's concept structure, but only when they are properly identified. This paper argues that similar breakthroughs occur frequently throughout training, but they are obscured by a loss metric that collapses all variation into a single scalar. To find these hidden transitions, we introduce POLCA, a method for decomposing changes in loss along arbitrary bases of the low-rank training subspace. We use our method to identify clusters of samples that share similar changes in loss during training, disaggregating the overall loss into that of smaller groups of conceptually similar data. We validate our method on synthetic arithmetic and English language modeling, showing that POLCA recovers clusters that represent interpretable breakthroughs in the model's capabilities. We demonstrate the promise of these hidden breakthroughs as a tool for unsupervised interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- The Mechanistic Emergence of Symbol Grounding in Language ModelsShuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang et al.ICML 2026 · 4 citations
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 3 citations
- In-Context AlgebraEric Todd, Jannik Brinkmann, Rohit Gandikota, David BauICLR 2026 · 3 citations
- ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model BehaviorFlorian Eichin, Yupei Du, Philipp Mondorf, Maria Matveev et al.ICML 2026 · 1 citation
- Understanding the Emergence of Seemingly Useless Features in Next-Token PredictorsMark Rofin, Jalal Naghiyev, Michael HahnICLR 2026
Builds on19
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 181 citations
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang et al.NeurIPS 2023 · 143 citations
Related papers
- Abrupt Learning in Transformers: A Case Study on Matrix CompletionPulkit Gopalani, Ekdeep Singh Lubana, Wei HuNeurIPS 2024 · 12 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
- Discovering and Steering Interpretable Concepts in Large Generative Music ModelsNikhil Singh, Manuel Cherep, Pattie MaesICLR 2026 · 17 citations
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised LearningXinting Huang, Michael HahnICLR 2026 · 7 citations
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri et al.EMNLP 2024 · 3 citations
