Interpretability Illusions in the Generalization of Simplified Models
Dan Friedman, Andrew Kyle Lampinen, Lucas Dixon, Danqi Chen, Asma Ghandeharioun
Abstract
A common method to study deep learning systems is to use simplified model representations--for example, using singular value decomposition to visualize the model's hidden states in a lower dimensional space. This approach assumes that the results of these simplifications are faithful to the original model. Here, we illustrate an important caveat to this assumption: even if the simplified representations can accurately approximate the full model on the training set, they may fail to accurately capture the model's behavior out of distribution. We illustrate this by training Transformer models on controlled datasets with systematic generalization splits, including the Dyck balanced-parenthesis languages and a code completion task. We simplify these models using tools like dimensionality reduction and clustering, and then explicitly test how these simplified proxies match the behavior of the original model. We find consistent generalization gaps: cases in which the simplified proxies are more faithful to the original model on the in-distribution evaluations and less faithful on various tests of systematic generalization. This includes cases where the original model generalizes systematically but the simplified proxies fail, and cases where the simplified proxies generalize better. Together, our results raise questions about the extent to which mechanistic interpretations derived using tools like SVD can reliably predict what a model will do in novel situations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi et al.NeurIPS 2024 · 196 citations
- Obfuscated Activations Bypass LLM Latent-Space DefensesLuke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov et al.ICLR 2026 · 28 citations
- Grokking Group Multiplication with CosetsDashiell Stander, Qinan Yu, Honglu Fan, Stella BidermanICML 2024 · 20 citations
- Predicting the Performance of Black-box Language Models with Follow-up QueriesDylan Sam, Marc Finzi, Zico KolterNeurIPS 2025 · 10 citations
- Certified Circuits: Stability Guarantees for Mechanistic CircuitsAlaa Anani, Tobias Lorenz, Bernt Schiele, Mario Fritz et al.ICML 2026 · 3 citations
Builds on18
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 307 citations
Related papers
- SVD as a Fast Interpretability Method for TransformersMin Xue, Artur AndrzejakICML 2026
- Quantifying and Optimizing Simplicity via Polynomial RepresentationsTianren Zhang, Xiangxin Li, Minghao Xiao, Guanyu Chen et al.ICML 2026
- Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann et al.ACL 2025 · 1 citation
- Transformers learn factored representationsAdam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow et al.ICML 2026 · 2 citations
- Lost in Latent Space: Examining failures of disentangled models at combinatorial generalisationMilton Llera Montero, Jeffrey S. Bowers, Rui Ponte Costa, Casimir J. H. Ludwig et al.NeurIPS 2022 · 29 citations
