Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Julian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt, Cameron Holmes, Robert West, Neel Nanda
摘要
Finetuning on narrow domains has become an essential tool to adapt Large Language Models (LLMs) to specific tasks and to create models with known unusual properties that are useful for safety research. Model diffing--the study of differences between base and finetuned models--is a promising approach for understanding how finetuning modifies neural networks. In this paper, we show that narrow finetuning creates easily readable biases in LLM activations that can be detected using simple model diffing tools, suggesting that the finetuning data is overrepresented in the model's activations. In particular, analyzing activation differences between base and finetuned models on the first few tokens of random text and steering with this difference allows us to recover the format and general content of the finetuning data. We call this the Activation Difference Lens (ADL). We demonstrate that these analyses significantly enhance an LLM-based interpretability agent's ability to identify subtle finetuning objectives through interaction with base and finetuned models. Our analysis spans synthetic document finetuning for false facts, emergent misalignment, subliminal learning, and taboo guessing game models across different architectures (Gemma, LLaMA, Qwen) and scales (1B to 32B parameters). Our work: (1) demonstrates that researchers should be aware that narrow finetuned models will represent their training data and objective very saliently, (2) warns AI safety and mechanistic interpretability researchers that these models might not be a realistic proxy for studying broader finetuning, despite current literature widely using them. While we show that mixing pretraining data into the finetuning corpus is enough to remove this bias, a deeper investigation is needed to understand the side effects of narrow finetuning and develop truly realistic case studies for model-diffing, safety and interpretability research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersAdam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente 等ICML 2026 · 被引用 42 次
- Introspection Adapters: Training LLMs to Report Their Learned BehaviorsKeshav Shenoy, Li Yang, Abhay Sheshadri, Jack Lindsey 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper12
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 被引用 146 次
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksSamyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick 等ICLR 2024 · 被引用 108 次
相关 Paper
- Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMsNikita Afonin, Nikita Andriyanov, Vahagn Hovhannisyan, Nikhil Bageshpura 等ACL 2026 · 被引用 12 次
- Understanding Emergent Misalignment via Feature Superposition GeometryGouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa 等ACL 2026 · 被引用 3 次
- Persona Features Control Emergent MisalignmentMiles Wang, Tom Dupré la Tour, Olivia Watkins, Aleksandar Makelov 等ICLR 2026 · 被引用 81 次
- Learning to Interpret Weight Differences in Language ModelsAvichal Goel, Yoon Kim, Nir N Shavit, Tony T. WangICLR 2026 · 被引用 10 次
- In-Training Defenses Against Emergent Misalignment in Language ModelsDavid Kaczér, Magnus Jørgenvåg, Clemens Vetter, Esha Afzal 等ICML 2026 · 被引用 13 次
