ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
Landon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas, Bin Yu, Kannan Ramchandran
Abstract
Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features. To identify these interactions, most existing approaches require enumerating all possible combinations of features up to a given order, causing them to scale poorly with the number of inputs n. Recently, Kang et al. (2025) proposed SPEX, an information-theoretic approach that uses interaction sparsity to scale to n ≈ 10 3 features. SPEX greatly improves upon prior methods but requires tens of thousands of model inferences, which can be prohibitive for large models. In this paper, we observe that LLM feature interactions are often hierarchical-higher-order interactions are accompanied by their lower-order subsets-which enables more efficient discovery. To exploit this hierarchy, we propose PROXYSPEX, an interaction attribution algorithm that first fits gradient boosted trees to masked LLM outputs and then extracts the important interactions. Experiments across four challenging high-dimensional datasets show that PROXYSPEX more faithfully reconstructs LLM outputs by 20% over marginal attribution approaches while using 10× fewer inferences than SPEX. By accounting for interactions, PROXYSPEX efficiently identifies the most influential features, providing a scalable approximation of their Shapley values. Further, we apply PROXYSPEX to two interpretability tasks. Data attribution, where we identify interactions among CIFAR-10 training samples that influence test predictions, and mechanistic interpretability, where we uncover interactions between attention heads, both within and across layers, on a question-answering task. The PROXYSPEX algorithm is available at https://github.com/mmschlk/shapiq.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74617ddf-b64d-44c5-88fb-e6cd6548b207Cited by top-tier papers6
- Regression-adjusted Monte Carlo Estimators for Shapley Values and Probabilistic ValuesR. Teal Witter, Yurong Liu, Christopher MuscoNeurIPS 2025 · 22 citations
- PolySHAP: Extending KernelSHAP with Interaction-Informed Polynomial RegressionFabian Fumagalli, R. Teal Witter, Christopher MuscoICLR 2026 · 7 citations
- Compressed Sensing for Capability Localization in Large Language ModelsAnna Bair, Yixuan Xu, Mingjie Sun, Zico KolterICML 2026 · 1 citation
- Evaluating and Explaining Prompt Sensitivity of LLMs Using InteractionsRuiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei et al.ICML 2026 · 1 citation
- SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) ModelsMingYu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim et al.ICML 2026
Builds on23
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- Scatterbrain: Unifying Sparse and Low-rank AttentionBeidi Chen, Tri Dao, Eric Winsor, Zhao Song et al.NeurIPS 2021 · 165 citations
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 118 citations
Related papers
- SPEX: Scaling Feature Interaction Explanations for LLMsJustin Singh Kang, Landon Butler, Abhineet Agarwal, Yigit Efe Erginbas et al.ICML 2025
- Multi-Level Explanations for Generative Language ModelsLucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt et al.ACL 2025 · 16 citations
- AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context AttributionFengyuan Liu, Nikhil Kandpal, Colin RaffelICLR 2025
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang et al.ICML 2026
- GiLOT: Interpreting Generative Language Models via Optimal TransportXuhong Li, Jiamin Chen, Yekun Chai, Haoyi XiongICML 2024 · 6 citations
