CIDR: A Cooperative Integrated Dynamic Refining Method for Minimal Feature Removal Problem
Qian Chen, Taolin Zhang, Dongyang Li, Xiaofeng He
Abstract
The minimal feature removal problem in the post-hoc explanation area aims to identify the minimal feature set (MFS). Prior studies using the greedy algorithm to calculate the minimal feature set lack the exploration of feature interactions under a monotonic assumption which cannot be satisfied in general scenarios. In order to address the above limitations, we propose a Cooperative Integrated Dynamic Refining method (CIDR) to efficiently discover minimal feature sets. Specifically, we design Cooperative Integrated Gradients (CIG) to detect interactions between features. By incorporating CIG and characteristics of the minimal feature set, we transform the minimal feature removal problem into a knapsack problem. Additionally, we devise an auxiliary Minimal Feature Refinement algorithm to determine the minimal feature set from numerous candidate sets. To the best of our knowledge, our work is the first to address the minimal feature removal problem in the field of natural language processing. Extensive experiments demonstrate that CIDR is capable of tracing representative minimal feature sets with improved interpretability across various models and datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 109 citations
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
Related papers
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- Axiomatic Aggregations of Abductive ExplanationsGagan Biradar, Yacine Izza, Elita A. Lobo, Vignesh Viswanathan et al.AAAI 2024 · 11 citations
- Joint Distribution–Informed Shapley Values for Sparse Counterfactual ExplanationsLei You, Yijun Bian, Lele CaoICLR 2026 · 3 citations
- Flexible Instance-Specific Rationalization of NLP ModelsGeorge Chrysostomou, Nikolaos AletrasAAAI 2022 · 17 citations
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
