ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
Landon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas, Bin Yu, Kannan Ramchandran
摘要
Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features. To identify these interactions, most existing approaches require enumerating all possible combinations of features up to a given order, causing them to scale poorly with the number of inputs n. Recently, Kang et al. (2025) proposed SPEX, an information-theoretic approach that uses interaction sparsity to scale to n ≈ 10 3 features. SPEX greatly improves upon prior methods but requires tens of thousands of model inferences, which can be prohibitive for large models. In this paper, we observe that LLM feature interactions are often hierarchical-higher-order interactions are accompanied by their lower-order subsets-which enables more efficient discovery. To exploit this hierarchy, we propose PROXYSPEX, an interaction attribution algorithm that first fits gradient boosted trees to masked LLM outputs and then extracts the important interactions. Experiments across four challenging high-dimensional datasets show that PROXYSPEX more faithfully reconstructs LLM outputs by 20% over marginal attribution approaches while using 10× fewer inferences than SPEX. By accounting for interactions, PROXYSPEX efficiently identifies the most influential features, providing a scalable approximation of their Shapley values. Further, we apply PROXYSPEX to two interpretability tasks. Data attribution, where we identify interactions among CIFAR-10 training samples that influence test predictions, and mechanistic interpretability, where we uncover interactions between attention heads, both within and across layers, on a question-answering task. The PROXYSPEX algorithm is available at https://github.com/mmschlk/shapiq.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Regression-adjusted Monte Carlo Estimators for Shapley Values and Probabilistic ValuesR. Teal Witter, Yurong Liu, Christopher MuscoNeurIPS 2025 · 被引用 22 次
- PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial RegressionFabian Fumagalli, R. Teal Witter, Christopher MuscoICLR 2026 · 被引用 7 次
- Compressed Sensing for Capability Localization in Large Language ModelsAnna Bair, Yixuan Xu, Mingjie Sun, Zico KolterICML 2026 · 被引用 1 次
- Evaluating and Explaining Prompt Sensitivity of LLMs Using InteractionsRuiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei 等ICML 2026 · 被引用 1 次
- SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) ModelsMingYu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim 等ICML 2026
它引用的顶会 Paper23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 被引用 199 次
- Scatterbrain: Unifying Sparse and Low-rank AttentionBeidi Chen, Tri Dao, Eric Winsor, Zhao Song 等NeurIPS 2021 · 被引用 165 次
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 被引用 118 次
相关 Paper
- SPEX: Scaling Feature Interaction Explanations for LLMsJustin Singh Kang, Landon Butler, Abhineet Agarwal, Yigit Efe Erginbas 等ICML 2025
- Multi-Level Explanations for Generative Language ModelsLucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt 等ACL 2025 · 被引用 16 次
- AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context AttributionFengyuan Liu, Nikhil Kandpal, Colin RaffelICLR 2025
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang 等ICML 2026
- GiLOT: Interpreting Generative Language Models via Optimal TransportXuhong Li, Jiamin Chen, Yekun Chai, Haoyi XiongICML 2024 · 被引用 6 次
