Tracking Equivalent Mechanistic Interpretations Across Neural Networks
Alan Sun, Mariya Toneva
摘要
Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that task. However, MI is difficult to scale and generalize. This stems in part from two key challenges: there is no precise notion of a valid interpretation; and, generating interpretations is often an ad hoc process. In this paper, we address these challenges by defining and studying the problem of interpretive equivalence: determining whether two different models share a common interpretation, without requiring an explicit description of what that interpretation is. At the core of our approach, we propose and formalize the principle that two interpretations of a model are equivalent if all of their possible implementations are also equivalent. We develop an algorithm to estimate interpretive equivalence and case study its use on Transformer-based models. To analyze our algorithm, we introduce necessary and sufficient conditions for interpretive equivalence based on models' representation similarity. We provide guarantees that simultaneously relate a model's algorithmic interpretations, circuits, and representations. Our framework lays a foundation for the development of more rigorous evaluation methods of MI and automated, generalizable interpretation discovery methods. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- Validating Mechanistic Interpretations: An Axiomatic ApproachNils Palumbo, Ravi Mangal, Zifan Wang, Saranya Vijayakumar 等ICML 2025
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?Denis Sutter, Julian Minder, Thomas Hofmann, Tiago PimentelNeurIPS 2025 · 被引用 30 次
- Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?Maxime Méloux, Silviu Maniu, François Portet, Maxime PeyrardICLR 2025
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre 等NeurIPS 2024 · 被引用 22 次
