Fresh in memory: Training-order recency is linearly encoded in language model activations
Dmitrii Krasheninnikov, Richard E. Turner, David Krueger
Abstract
We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentially fine-tuning Llama-3.2-1B on six disjoint but otherwise similar datasets about named entities. We find that the average activations of test samples corresponding to the six training datasets encode the training order: when projected into a 2D subspace, these centroids are arranged exactly in the order of training and lie on a straight line. Further, we show that linear probes can accurately (∼90%) distinguish "early" vs. "late" entities, generalizing to entities unseen during the probes' own training. The model can also be fine-tuned to explicitly report an unseen entity's training stage (∼80% accuracy). Interestingly, the training-order encoding does not seem attributable to simple differences in activation magnitudes, losses, or model confidence. Our paper demonstrates that models are capable of differentiating information by its acquisition time, and carries significant implications for how they might manage conflicting data and respond to knowledge modifications. INTRODUCTION Average centroid difference (stage 1 -stage 6) Top PC orthogonal to x-axis Activation centroids (averages) for the six test datasets, across four independent training runs Synth -who Synth -stand for Natural -name Natural -meaning Actual training order D1 D2 D3 D4 D5 D6
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5111600e-1245-4198-87ff-2f33bb4dfdbbCited by top-tier papers1
Ask how each one uses itBuilds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Manipulating SGD with Data Ordering AttacksIlia Shumailov, Zakhar Shumaylov, Dmitry Kazhdan, Yiren Zhao et al.NeurIPS 2021 · 125 citations
- Fortuitous Forgetting in Connectionist NetworksHattie Zhou, Ankit Vani, Hugo Larochelle, Aaron C. CourvilleICLR 2022 · 50 citations
- Probing Representation Forgetting in Supervised and Unsupervised Continual LearningMohammadReza Davari, Nader Asadi, Sudhir P. Mudur, Rahaf Aljundi et al.CVPR 2022 · 48 citations
- Implicit meta-learning may lead language models to trust more reliable sourcesDmitrii Krasheninnikov, Egor Krasheninnikov, Bruno Kacper Mlodozeniec, Tegan Maharaj et al.ICML 2024 · 8 citations
Related papers
- Understanding Data Temporality Impact on Large Language Models Pre-trainingRomain Fabre, Hippolyte Pilchen, Franck SIGNE TALLA, Patrick Perez et al.ICML 2026
- Word Order Does Matter and Shuffled Language Models Know ItMostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders SøgaardACL 2022
- Extractive Structures Learned in Pretraining Enable Generalization on Finetuned FactsJiahai Feng, Stuart Russell, Jacob SteinhardtICML 2025
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau et al.EMNLP 2021 · 177 citations
