The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable Importance
Jon Donnelly, Srikar Katta, Cynthia Rudin, Edward P. Browne
摘要
Quantifying variable importance is essential for answering high-stakes questions in fields like genetics, public policy, and medicine. Current methods generally calculate variable importance for a given model trained on a given dataset. However, for a given dataset, there may be many models that explain the target outcome equally well; without accounting for all possible explanations, different researchers may arrive at many conflicting yet equally valid conclusions given the same data. Additionally, even when accounting for all possible explanations for a given dataset, these insights may not generalize because not all good explanations are stable across reasonable data perturbations. We propose a new variable importance framework that quantifies the importance of a variable across the set of all good models and is stable across the data distribution. Our framework is extremely flexible and can be integrated with most existing model classes and global variable importance metrics. We demonstrate through experiments that our framework recovers variable importance rankings for complex simulation setups where other methods fail. Further, we show that our framework accurately estimates the true importance of a variable for the underlying data distribution. We provide theoretical guarantees on the consistency and finite sample error rates for our estimator. Finally, we demonstrate its utility with a real-world case study exploring which genes are important for predicting HIV load in persons with HIV, highlighting an important gene that has not previously been studied in connection with HIV. Code is available here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr 等NeurIPS 2024 · 被引用 10 次
- SORTeD Rashomon Sets of Sparse Decision Trees: Anytime EnumerationElif Arslan, Jacobus G. M. van der Linden, Serge P. Hoogendoorn, Marco Rinaldi 等NeurIPS 2025 · 被引用 8 次
- ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon SetsBohdan Turbal, Iryna Voitsitska, Lesia SemenovaNeurIPS 2025 · 被引用 6 次
- Fixed Aggregation Features Can Rival GNNsCelia Rubio-Madrigal, Rebekka BurkholzICML 2026 · 被引用 2 次
它引用的顶会 Paper5
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Generalized and Scalable Optimal Sparse Decision TreesJimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin 等ICML 2020 · 被引用 174 次
- Exploring the Whole Rashomon Set of Sparse Decision TreesRui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi 等NeurIPS 2022 · 被引用 117 次
- Efficient nonparametric statistical inference on population feature importance using Shapley valuesBrian D. Williamson, Jean FengICML 2020 · 被引用 86 次
- Model Class Reliance for Random ForestsGavin Smith, Roberto Mansilla, James GouldingNeurIPS 2020 · 被引用 33 次
相关 Paper
- Covered Information Disentanglement: Model Transparency via Unbiased Permutation ImportanceJoão P. B. Pereira, Erik S. G. Stroes, Aeilko H. Zwinderman, Evgeni LevinAAAI 2022 · 被引用 17 次
- Measuring Variable Importance in Heterogeneous Treatment Effects with ConfidenceJoseph Paillard, Angel David Reyero Lobo, Vitaliy Kolodyazhniy, Bertrand Thirion 等ICML 2025
- Aggregate Models, Not Explanations: Improving Feature Importance EstimationJoseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion BertrandICML 2026 · 被引用 1 次
- Statistically Valid Variable Importance Assessment through Conditional PermutationsAhmad Chamma, Denis A. Engemann, Bertrand ThirionNeurIPS 2023 · 被引用 23 次
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
