Selective Ensembles for Consistent Predictions
Emily Black, Klas Leino, Matt Fredrikson
Abstract
Recent work has shown that models trained to the same objective, and which achieve similar measures of accuracy on consistent test data, may nonetheless behave very differently on individual predictions. This inconsistency is undesirable in high-stakes contexts, such as medical diagnosis and finance. We show that this inconsistent behavior extends beyond predictions to feature attributions, which may likewise have negative implications for the intelligibility of a model, and one's ability to find recourse for subjects. We then introduce selective ensembles to mitigate such inconsistencies by applying hypothesis testing to the predictions of a set of models trained using randomly-selected starting conditions; importantly, selective ensembles can abstain in cases where a consistent outcome cannot be achieved up to a specified confidence level. We prove that that prediction disagreement between selective ensembles is bounded, and empirically demonstrate that selective ensembles achieve consistent predictions and feature attributions while maintaining low abstention rates. On several benchmark datasets, selective ensembles reach zero inconsistently predicted points, with abstention rates as low 1.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27d8f7e7-7ce3-4911-a624-adf035113e3aCited by top-tier papers11
- Predictive Multiplicity in Probabilistic ClassificationJamelle Watson-Daniels, David C. Parkes, Berk UstunAAAI 2023 · 58 citations
- Implications of Model Indeterminacy for Explanations of Automated DecisionsMarc-Etienne Brunet, Ashton Anderson, Richard S. ZemelNeurIPS 2022 · 22 citations
- Individual Arbitrariness and Group FairnessCarol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flávio P. CalmonNeurIPS 2023 · 16 citations
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Monoculture or Multiplicity: Which Is It?Mila Gorecki, Moritz HardtNeurIPS 2025 · 6 citations
Builds on9
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
- Diversity With Cooperation: Ensemble Methods for Few-Shot ClassificationNikita Dvornik, Julien Mairal, Cordelia SchmidICCV 2019 · 210 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 197 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
Related papers
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 6 citations
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen et al.AAAI 2026
- Training Private Models That Know What They Don't KnowStephan Rabanser, Anvith Thudi, Abhradeep Guha Thakurta, Krishnamurthy Dvijotham et al.NeurIPS 2023 · 10 citations
- Rejectors in the Wild: Deployment Barriers for LLM RejectorsErik Schönwälder, Claudio Hartmann, Wolfgang LehnerKDD 2026
- Improving Selective Visual Question Answering by Learning from Your PeersCorentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam et al.CVPR 2023
