Lune

NeurIPS2024Top-tier venue

Optimal ablation for interpretability

Maximilian Li, Lucas Janson

2024Year
32Citations
6Top-tier citations

Abstract

Interpretability studies often involve tracing the flow of information through machine learning models to identify specific model components that perform relevant computations for tasks of interest. Prior work quantifies the importance of a model component on a particular task by measuring the impact of performing ablation on that component, or simulating model inference with the component disabled. We propose a new method, optimal ablation (OA), and show that OA-based component importance has theoretical and empirical advantages over measuring importance via other ablation methods. We also show that OA-based component importance can benefit several downstream interpretability tasks, including circuit discovery, localization of factual recall, and latent prediction.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4b5f64a5-df79-4929-9cab-b13bc01a0356

Cited by top-tier papers6

Ask how each one uses it

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines