Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
Julian Minder, Clément Dumas, Caden Juang, Bilal Chughtai, Neel Nanda
摘要
Model diffing is the study of how fine-tuning changes a model's representations and internal algorithms. Many behaviors of interest are introduced during fine-tuning, and model diffing offers a promising lens to interpret such behaviors. Crosscoders are a recent model diffing method that learns a shared dictionary of interpretable concepts represented as latent directions in both the base and fine-tuned models, allowing us to track how concepts shift or emerge during fine-tuning. Notably, prior work has observed concepts with no direction in the base model, and it was hypothesized that these model-specific latents were concepts introduced during fine-tuning. However, we identify two issues which stem from the crosscoders L1 training loss that can misattribute concepts as unique to the fine-tuned model, when they really exist in both models. We develop Latent Scaling to flag these issues by more accurately measuring each latent's presence across models. In experiments comparing Gemma 2 2B base and chat models, we observe that the standard crosscoder suffers heavily from these issues. Building on these insights, we train a crosscoder with BatchTopK loss and show that it substantially mitigates these issues, finding more genuinely chat-specific and highly interpretable concepts. We recommend practitioners adopt similar techniques. Using the BatchTopK crosscoder, we successfully identify a set of chat-specific latents that are both interpretable and causally effective, representing concepts such as false information and personal question, along with multiple refusal-related latents that show nuanced preferences for different refusal triggers. Overall, our work advances best practices for the crosscoder-based methodology for model diffing and demonstrates that it can provide concrete insights into how chat-tuning modifies model behavior. 1 * Equal contribution. Order randomized. 1 We open-source our code, training library, models, wandb runs and a demo notebook to explore latents.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim 等ICML 2026 · 被引用 22 次
- Base Models Know How to Reason, Thinking Models Learn WhenConstantin Venhoff, Iván Arcuschin, Phil Torr, Arthur Conmy 等ICML 2026 · 被引用 20 次
- Learning to Interpret Weight Differences in Language ModelsAvichal Goel, Yoon Kim, Nir N Shavit, Tony T. WangICLR 2026 · 被引用 10 次
- Evolution of Concepts in Language Model Pre-TrainingXuyang Ge, Wentao Shu, Jiaxing Wu, Yunhua Zhou 等ICLR 2026 · 被引用 8 次
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 被引用 3 次
它引用的顶会 Paper22
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu 等NeurIPS 2025 · 被引用 181 次
- A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityAndrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg 等ICML 2024 · 被引用 177 次
相关 Paper
- Efficient Transcoder Adaptation for Fine-Tuned Models: Revealing Medical Reasoning Mechanisms in Large Language ModelsZhouxing Tan, Hanlin Xue, Yulong Wan, Ruochong Xiong 等AAAI 2026
- Narrow Finetuning Leaves Clearly Readable Traces in Activation DifferencesJulian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt 等ICLR 2026 · 被引用 29 次
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 被引用 8 次
- Persona Features Control Emergent MisalignmentMiles Wang, Tom Dupré la Tour, Olivia Watkins, Aleksandar Makelov 等ICLR 2026 · 被引用 81 次
- REVIVING YOUR MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model DiffingAly M. Kassem, Zhuan Shi, Negar Rostamzadeh, Golnoosh FarnadiEMNLP 2025
