Structural Inference: Interpreting Small Language Models with Susceptibilities
Garrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak, Daniel Murfet
Abstract
We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b726ce50-b5df-4b24-9c51-6b5cb44753a3Cited by top-tier papers4
- Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM UnitsJianhui Chen, Yuzhang Luo, Liangming PanICML 2026 · 6 citations
- Patterning: The Dual of InterpretabilityGeorge Wang, Daniel MurfetICML 2026 · 5 citations
- Towards Spectroscopy: Susceptibility Clusters in Language ModelsAndrew Gordon, Garrett Baker, George Wang, William Snell et al.ICML 2026 · 1 citation
- Mechanistic Anomaly Detection via Functional AttributionHugo Lyons Keenan, Christopher Leckie, Sarah ErfaniICML 2026
Builds on10
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 71 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
Related papers
- Measuring the Mixing of Contextual Information in the TransformerJavier Ferrando, Gerard I. Gállego, Marta R. Costa-jussàEMNLP 2022 · 18 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- SVD as a Fast Interpretability Method for TransformersMin Xue, Artur AndrzejakICML 2026
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov et al.ICLR 2026 · 8 citations
- DePass: Unified Feature Attributing by Simple Decomposed Forward PassXiangyu Hong, Che Jiang, Kai Tian, Biqing Qi et al.NeurIPS 2025 · 4 citations
