How does This Interaction Affect Me? Interpretable Attribution for Feature Interactions
Michael Tsang, Sirisha Rambhatla, Yan Liu
Abstract
Machine learning transparency calls for interpretable explanations of how inputs relate to predictions. Feature attribution is a way to analyze the impact of features on predictions. Feature interactions are the contextual dependence between features that jointly impact predictions. There are a number of methods that extract feature interactions in prediction models; however, the methods that assign attributions to interactions are either uninterpretable, model-specific, or non-axiomatic. We propose an interaction attribution and detection framework called Archipelago which addresses these problems and is also scalable in real-world settings. Our experiments on standard annotation labels indicate our approach provides significantly more interpretable explanations than comparable methods, which is important for analyzing the impact of interactions on predictions. We also provide accompanying visualizations of our approach that give new insights into deep neural networks. To this end, we propose a novel framework called Archipelago, which consists of an interaction attribution method, ArchAttribute, and a corresponding interaction detector, ArchDetect, to address the challenges of being interpretable, axiomatic, and scalable. Archipelago is named after its ability to provide explanations by isolating feature interactions, or feature "islands". The inputs to Archipelago are a black-box model f and data instance x , and its outputs are a set of interactions and individual features I as well as an attribution score φ(I) for each of the feature sets I. ArchAttribute satisfies attribution axioms by making relatively mild assumptions: a) disjointness of interaction sets, which is easily obtainable, and b) the availability of a generalized additive Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f85454a-5fe4-444b-a403-728afac03febCited by top-tier papers26
- SHAP-IQ: Unified Approximation of any-order Shapley InteractionsFabian Fumagalli, Maximilian Muschalik, Patrick Kolpaczki, Eyke Hüllermeier et al.NeurIPS 2023 · 80 citations
- Sparse Interaction Additive Networks via Feature Interaction Detection and Sparse SelectionJames Enouen, Yan LiuNeurIPS 2022 · 38 citations
- Beyond TreeSHAP: Efficient Computation of Any-Order Shapley Interactions for Tree EnsemblesMaximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke HüllermeierAAAI 2024 · 35 citations
- Explanations of Black-Box Models based on Directional Feature InteractionsAria Masoomi, Davin Hill, Zhonghui Xu, Craig P. Hersh et al.ICLR 2022 · 26 citations
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
Builds on3
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction DetectionMichael Tsang, Dehua Cheng, Hanpeng Liu, Xue Feng et al.ICLR 2020 · 71 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
Related papers
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 27 citations
- A Unifying Framework to the Analysis of Interaction Methods using Synergy FunctionsDaniel Lundström, Meisam RazaviyaynICML 2023 · 4 citations
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
