Selective Ensembles for Consistent Predictions
Emily Black, Klas Leino, Matt Fredrikson
摘要
Recent work has shown that models trained to the same objective, and which achieve similar measures of accuracy on consistent test data, may nonetheless behave very differently on individual predictions. This inconsistency is undesirable in high-stakes contexts, such as medical diagnosis and finance. We show that this inconsistent behavior extends beyond predictions to feature attributions, which may likewise have negative implications for the intelligibility of a model, and one's ability to find recourse for subjects. We then introduce selective ensembles to mitigate such inconsistencies by applying hypothesis testing to the predictions of a set of models trained using randomly-selected starting conditions; importantly, selective ensembles can abstain in cases where a consistent outcome cannot be achieved up to a specified confidence level. We prove that that prediction disagreement between selective ensembles is bounded, and empirically demonstrate that selective ensembles achieve consistent predictions and feature attributions while maintaining low abstention rates. On several benchmark datasets, selective ensembles reach zero inconsistently predicted points, with abstention rates as low 1.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Predictive Multiplicity in Probabilistic ClassificationJamelle Watson-Daniels, David C. Parkes, Berk UstunAAAI 2023 · 被引用 58 次
- Implications of Model Indeterminacy for Explanations of Automated DecisionsMarc-Etienne Brunet, Ashton Anderson, Richard S. ZemelNeurIPS 2022 · 被引用 22 次
- Individual Arbitrariness and Group FairnessCarol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flávio P. CalmonNeurIPS 2023 · 被引用 16 次
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Monoculture or Multiplicity: Which Is It?Mila Gorecki, Moritz HardtNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper9
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 被引用 774 次
- Diversity With Cooperation: Ensemble Methods for Few-Shot ClassificationNikita Dvornik, Julien Mairal, Cordelia SchmidICCV 2019 · 被引用 210 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
相关 Paper
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 被引用 6 次
- DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image DiagnosticsHong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen 等AAAI 2026
- Training Private Models That Know What They Don't KnowStephan Rabanser, Anvith Thudi, Abhradeep Guha Thakurta, Krishnamurthy Dvijotham 等NeurIPS 2023 · 被引用 10 次
- Rejectors in the Wild: Deployment Barriers for LLM RejectorsErik Schönwälder, Claudio Hartmann, Wolfgang LehnerKDD 2026
- Improving Selective Visual Question Answering by Learning from Your PeersCorentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam 等CVPR 2023
