Model Class Reliance for Random Forests
Gavin Smith, Roberto Mansilla, James Goulding
Abstract
Variable Importance (VI) has traditionally been cast as the process of estimating each variable's contribution to a predictive model's overall performance. Analysis of a single model instance, however, guarantees no insight into a variables relevance to underlying generative processes. Recent research has sought to address this concern via analysis of Rashomon sets -sets of alternative model instances that exhibit equivalent predictive performance to some reference model, but which take different functional forms. Measures such as Model Class Reliance (MCR) have been proposed, that are computed against Rashomon sets, in order to ascertain how much a variable must be relied on to make robust predictions, or whether alternatives exist. If MCR range is tight, we have no choice but to use a variable; if range is high then there exists competing, perhaps fairer models, that provide alternative explanations of the phenomena being examined. Applications are wide, from enabling construction of 'fairer' models in areas such as recidivism to health analytics and ethical marketing. Tractable estimation of MCR for nonlinear models is currently restricted to Kernel Regression under squared loss [7] . In this paper we introduce a new technique that extends computation of Model Class Reliance (MCR) to Random Forest classifiers and regressors. The proposed approach addresses a number of open research questions, and in contrast to prior Kernel SVM MCR estimation, runs in linearithmic rather than polynomial time. Taking a fundamentally different approach to previous work, we provide a solution for this important model class, identifying situations where irrelevant covariates do not improve predictions. Recently, [7] formally defined this issue -the multiplicity of models with the same prediction performance based on different underpining predicting factors -with regard to model classes. Given 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36a4ad58-a037-4631-a0d0-64fc32cc7ac4Cited by top-tier papers7
- Exploring the Whole Rashomon Set of Sparse Decision TreesRui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi et al.NeurIPS 2022 · 117 citations
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 41 citations
- The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable ImportanceJon Donnelly, Srikar Katta, Cynthia Rudin, Edward P. BrowneNeurIPS 2023 · 41 citations
- Exploring and Interacting with the Set of Good Sparse Generalized Additive ModelsChudi Zhong, Zhi Chen, Jiachang Liu, Margo I. Seltzer et al.NeurIPS 2023 · 39 citations
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr et al.NeurIPS 2024 · 10 citations
Related papers
- Exploring the cloud of feature interaction scores in a Rashomon setSichao Li, Rong Wang, Quanling Deng, Amanda S. BarnardICLR 2024 · 9 citations
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 98 citations
- Rashomon Capacity: A Metric for Predictive Multiplicity in ClassificationHsiang Hsu, Flávio P. CalmonNeurIPS 2022 · 65 citations
- Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity EstimationHsiang Hsu, Guihong Li, Shaohan Hu, Chun-Fu ChenICLR 2024 · 18 citations
