Optimal ablation for interpretability
Maximilian Li, Lucas Janson
Abstract
Interpretability studies often involve tracing the flow of information through machine learning models to identify specific model components that perform relevant computations for tasks of interest. Prior work quantifies the importance of a model component on a particular task by measuring the impact of performing ablation on that component, or simulating model inference with the component disabled. We propose a new method, optimal ablation (OA), and show that OA-based component importance has theoretical and empirical advantages over measuring importance via other ablation methods. We also show that OA-based component importance can benefit several downstream interpretability tasks, including circuit discovery, localization of factual recall, and latent prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b5f64a5-df79-4929-9cab-b13bc01a0356Cited by top-tier papers6
- Discovering Interpretable Algorithms by Decompiling Transformers to RASPXinting Huang, Aleksandra Bakalova, Satwik Bhattamishra, William Merrill et al.ICML 2026 · 3 citations
- Language Model Circuits Are Sparse in the Neuron BasisAryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah SchwettmannICML 2026
- MIB: A Mechanistic Interpretability BenchmarkAaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad et al.ICML 2025
- All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other TokensSiddarth Mamidanna, Daking Rai, Ziyu Yao, Yilun ZhouEMNLP 2025
- GKnow: Measuring the Entanglement of Gender Bias and Factual GenderLeonor Veloso, Hinrich SchützeACL 2026
Builds on35
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
Related papers
- Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster InferenceJorge García-Carrasco, Alejandro Maté, Juan TrujilloAAAI 2025 · 3 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
- Explainability as statistical inferenceHugo Henri Joseph Senetaire, Damien Garreau, Jes Frellsen, Pierre-Alexandre MatteiICML 2023 · 4 citations
- Decomposing and Editing Predictions by Modeling Model ComputationHarshay Shah, Andrew Ilyas, Aleksander MadryICML 2024 · 25 citations
- Leveraging Predictive Equivalence in Decision TreesHayden McTavish, Zachery Boner, Jon Donnelly, Margo I. Seltzer et al.ICML 2025
