Evolution of Concepts in Language Model Pre-Training
Xuyang Ge, Wentao Shu, Jiaxing Wu, Yunhua Zhou, Zhengfu He, Xipeng Qiu
Abstract
Language models obtain extensive capabilities through pre-training. However, the pre-training dynamics remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics. Our code is available at https://github.com/OpenMOSS/Language-Model-SAEs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2518f36f-032e-4b85-ab59-ef66015cae1eCited by top-tier papers2
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen et al.ICLR 2026 · 6 citations
- Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM UnitsJianhui Chen, Yuzhang Luo, Liangming PanICML 2026 · 6 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
Related papers
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 3 citations
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for InterpretabilityUsha Bhalla, Alex Oesterling, Claudio Mayrink Verdun, Himabindu Lakkaraju et al.ICLR 2026 · 18 citations
- Analyze Feature Flow to Enhance Interpretation and Steering in Language ModelsDaniil Laptev, Nikita Balagansky, Yaroslav Aksenov, Daniil GavrilovICML 2025
- Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsAdam Karvonen, Benjamin Wright, Can Rager, Rico Angell et al.NeurIPS 2024 · 66 citations
- Exploring the Benefit of Activation Sparsity in Pre-trainingZhengyan Zhang, Chaojun Xiao, Qiujieli Qin, Yankai Lin et al.ICML 2024 · 6 citations
