The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel
摘要
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there exists a function which allows us to map between them. Notably, most interpretability papers implement these maps as linear functions, motivated by the linear representation hypothesis: the idea that features are encoded linearly in a model's representations. However, this linearity constraint is not required by the definition of causal abstraction. In this work, we critically examine the concept of causal abstraction by considering arbitrarily powerful alignment maps. In particular, we prove that under reasonable assumptions, any neural network can be mapped to any algorithm, rendering this unrestricted notion of causal abstraction trivial and uninformative. We complement these theoretical findings with empirical evidence, demonstrating that it is possible to perfectly map models to algorithms even when these models are incapable of solving the actual task; e.g., on an experiment using randomly initialised language models, our alignment maps reach 100% interchange-intervention accuracy on the indirect object identification task. This raises the non-linear representation dilemma: if we lift the linearity constraint imposed to alignment maps in causal abstraction analyses, we are left with no principled way to balance the inherent trade-off between these maps'complexity and accuracy. Together, these results suggest an answer to our title's question: causal abstraction is not enough for mechanistic interpretability, as it becomes vacuous without assumptions about how models encode information. Studying the connection between this information-encoding assumption and causal abstraction should lead to exciting future work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Language Models are Injective and Hence InvertibleGiorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli 等ICLR 2026 · 被引用 36 次
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa 等ICLR 2026 · 被引用 19 次
- Addressing divergent representations from causal interventions on neural networksSatchel Grant, Simon Jerome Han, Alexa R. Tartaglini, Christopher PottsICLR 2026 · 被引用 7 次
- Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative ReviewingMichael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi, Philip Quirke 等ACL 2026 · 被引用 1 次
- Operationalising the Superficial Alignment Hypothesis via Task ComplexityTomás Vergara Browne, Darshan Patil, Ivan Titov, Siva Reddy 等ICML 2026
它引用的顶会 Paper17
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts 等NeurIPS 2023 · 被引用 146 次
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner 等ICML 2022 · 被引用 104 次
相关 Paper
- Linear Causal Representation Learning by Topological Ordering, Pruning, and DisentanglementHao Chen, Lin Liu, Yuguang WangICML 2026
- Tracking Equivalent Mechanistic Interpretations Across Neural NetworksAlan Sun, Mariya TonevaICLR 2026 · 被引用 1 次
- Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?Maxime Méloux, Silviu Maniu, François Portet, Maxime PeyrardICLR 2025
- From Neurons to Neutrons: A Case Study in InterpretabilityOuail Kitouni, Niklas Nolte, Víctor Samuel Pérez-Díaz, Sokratis Trifinopoulos 等ICML 2024 · 被引用 4 次
- Validating Mechanistic Interpretations: An Axiomatic ApproachNils Palumbo, Ravi Mangal, Zifan Wang, Saranya Vijayakumar 等ICML 2025
