Interpretability Illusions in the Generalization of Simplified Models
Dan Friedman, Andrew Kyle Lampinen, Lucas Dixon, Danqi Chen, Asma Ghandeharioun
摘要
A common method to study deep learning systems is to use simplified model representations--for example, using singular value decomposition to visualize the model's hidden states in a lower dimensional space. This approach assumes that the results of these simplifications are faithful to the original model. Here, we illustrate an important caveat to this assumption: even if the simplified representations can accurately approximate the full model on the training set, they may fail to accurately capture the model's behavior out of distribution. We illustrate this by training Transformer models on controlled datasets with systematic generalization splits, including the Dyck balanced-parenthesis languages and a code completion task. We simplify these models using tools like dimensionality reduction and clustering, and then explicitly test how these simplified proxies match the behavior of the original model. We find consistent generalization gaps: cases in which the simplified proxies are more faithful to the original model on the in-distribution evaluations and less faithful on various tests of systematic generalization. This includes cases where the original model generalizes systematically but the simplified proxies fail, and cases where the simplified proxies generalize better. Together, our results raise questions about the extent to which mechanistic interpretations derived using tools like SVD can reliably predict what a model will do in novel situations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi 等NeurIPS 2024 · 被引用 196 次
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov 等ICLR 2026 · 被引用 28 次
- Grokking Group Multiplication with CosetsDashiell Stander, Qinan Yu, Honglu Fan, Stella BidermanICML 2024 · 被引用 20 次
- Predicting the Performance of Black-box Language Models with Follow-up QueriesDylan Sam, Marc Finzi, Zico KolterNeurIPS 2025 · 被引用 10 次
- Certified Circuits: Stability Guarantees for Mechanistic CircuitsAlaa Anani, Tobias Lorenz, Bernt Schiele, Mario Fritz 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper18
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 被引用 307 次
相关 Paper
- SVD as a Fast Interpretability Method for TransformersMin Xue, Artur AndrzejakICML 2026
- Quantifying and Optimizing Simplicity via Polynomial RepresentationsTianren Zhang, Xiangxin Li, Minghao Xiao, Guanyu Chen 等ICML 2026
- Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann 等ACL 2025 · 被引用 1 次
- Transformers learn factored representationsAdam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow 等ICML 2026 · 被引用 2 次
- Lost in Latent Space: Examining failures of disentangled models at combinatorial generalisationMilton Llera Montero, Jeffrey S. Bowers, Rui Ponte Costa, Casimir J. H. Ludwig 等NeurIPS 2022 · 被引用 29 次
