Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang, William Snell, Stan van Wingerden, Daniel Murfet
摘要
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token in context , we measure the model's response via susceptibilities , which are covariances between component-level observables and the perturbation computed over a localized Gibbs posterior via stochastic gradient Langevin dynamics (SGLD). Theoretically, we show that susceptibilities decompose as a sum over modes of the data distribution, explaining why tokens that follow their contexts ``for similar reasons'' cluster together in susceptibility space. Empirically, we apply this methodology to Pythia-14M, developing a conductance-based clustering algorithm that identifies 510 interpretable clusters ranging from grammatical patterns to code structure to mathematical notation. Comparing to sparse autoencoders, 50% of our clusters match SAE features, validating that both methods recover similar structure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani 等ICLR 2022 · 被引用 74 次
- Bayesian Influence Functions for Hessian-Free Data AttributionPhilipp Alexander Kreer, Wilson Wu, Maxwell Adam, Zach Furman 等ICLR 2026 · 被引用 14 次
相关 Paper
- Structural Inference: Interpreting Small Language Models with SusceptibilitiesGarrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak 等ICLR 2026 · 被引用 11 次
- Towards Universality: Studying Mechanistic Similarity Across Language Model ArchitecturesJunxuan Wang, Xuyang Ge, Wentao Shu, Qiong Tang 等ICLR 2025
- The Deleuzian Representation HypothesisClément Cornet, Romaric Besançon, Hervé Le BorgneICLR 2026
- Discovering and Steering Interpretable Concepts in Large Generative Music ModelsNikhil Singh, Manuel Cherep, Pattie MaesICLR 2026 · 被引用 17 次
- I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?Yuhang Liu, Dong Gong, Yichao Cai, Erdun Gao 等ICLR 2026 · 被引用 17 次
