Variable Importance in High-Dimensional Settings Requires Grouping
Ahmad Chamma, Bertrand Thirion, Denis A. Engemann
摘要
Explaining the decision process of machine learning algorithms is nowadays crucial for both a model's performance enhancement and human comprehension. This can be achieved by assessing the variable importance of single variables, even for high-capacity non-linear methods, e.g. Deep Neural Networks (DNNs). While only removalbased approaches, such as Permutation Importance (PI), can bring statistical validity, they return misleading results when variables are correlated. Conditional Permutation Importance (CPI) bypasses PI's limitations in such cases. However, in high-dimensional settings, where high correlations between the variables cancel their conditional importance, the use of CPI as well as other methods leads to unreliable results, besides prohibitive computation costs. Grouping variables statistically via clustering or some prior knowledge gains some power back and leads to better interpretations. In this work, we introduce BCPI (Block-Based Conditional Permutation Importance), a new generic framework for variable importance computation with statistical guarantees handling both single and group cases. Furthermore, as handling groups with high cardinality (such as a set of observations of a given modality) are both time-consuming and resource-intensive, we also introduce a new stacking approach extending the DNN architecture with sub-linear layers adapted to the group structure. We show that the ensuing approach extended with stacking controls the type-I error even with highly-correlated groups and shows top accuracy across benchmarks. Furthermore, we perform a real-world data analysis in a large-scale medical dataset where we aim to show the consistency between our results and the literature for a biomarker prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Flow-Disentangled Feature ImportanceXingshu Chen, Yifeng Guo, Jin-Hong DuICLR 2026
- Measuring Variable Importance in Heterogeneous Treatment Effects with ConfidenceJoseph Paillard, Angel David Reyero Lobo, Vitaliy Kolodyazhniy, Bertrand Thirion 等ICML 2025
它引用的顶会 Paper3
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- Statistically Valid Variable Importance Assessment through Conditional PermutationsAhmad Chamma, Denis A. Engemann, Bertrand ThirionNeurIPS 2023 · 被引用 23 次
- Lazy Estimation of Variable Importance for Large Neural NetworksYue Gao, Abby Stevens, Garvesh Raskutti, Rebecca WillettICML 2022 · 被引用 7 次
相关 Paper
- Testing Conditional Mean Independence Using Generative Neural NetworksYi Zhang, Linjun Huang, Yun Yang, Xiaofeng ShaoICML 2025
- Covered Information Disentanglement: Model Transparency via Unbiased Permutation ImportanceJoão P. B. Pereira, Erik S. G. Stroes, Aeilko H. Zwinderman, Evgeni LevinAAAI 2022 · 被引用 17 次
- Aggregate Models, Not Explanations: Improving Feature Importance EstimationJoseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion BertrandICML 2026 · 被引用 1 次
- Permutation-Based Hypothesis Testing for Neural NetworksFrancesca Mandel, Ian BarnettAAAI 2024 · 被引用 6 次
- Explainability as statistical inferenceHugo Henri Joseph Senetaire, Damien Garreau, Jes Frellsen, Pierre-Alexandre MatteiICML 2023 · 被引用 4 次
