DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image Diagnostics
Hong Hai Nguyen, Duong Bach, Nam Phan, Cuong V. Nguyen, Cuong Do
Abstract
For artificial intelligence to be safely deployed in high-risk domains, it must reliably know its limits. Selective prediction, or learning with a reject option, addresses this by enabling a model to abstain from prediction on inputs it deems unreliable, deferring them to a human expert. While deep ensembles have emerged as a leading approach for uncertainty estimation, their potential is often squandered by rejection methods that rely on static thresholds applied to the mean prediction. In this paper, we propose to learn a dynamic rejection policy directly from the rich behavioral signals of the ensemble itself. Our framework, DEGRE (Dynamic Ensembles Gating for REjection), is a novel meta-learning approach that trains a lightweight gating network on the ensemble’s consensus confidence and its internal disagreement (variance)— to explicitly discriminate between correct and incorrect predictions. Through rigorous evaluation across twelve diverse medical imaging benchmarks (MRI, X-ray, CT), DEGRE significantly advances selective prediction, achieving an average risk-coverage (AURC) reduction of 68.2% compared to the standard ensemble baseline. By providing a more reliable method for a model to recognize its own limitations, this learned, adaptive rejection mechanism paves the way for safer and more responsible integration of AI into critical clinical workflows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 250 citations
- Role of Human-AI Interaction in Selective PredictionElizabeth Bondi, Raphael Koster, Hannah Sheahan, Martin J. Chadwick et al.AAAI 2022 · 46 citations
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 11 citations
- Generative Medical SegmentationJiayu Huo, Xi Ouyang, Sébastien Ourselin, Rachel SparksAAAI 2025 · 8 citations
Related papers
- Top-Ambiguity Samples Matter: Understanding Why Deep Ensemble Works in Selective ClassificationQiang Ding, Yixuan Cao, Ping LuoNeurIPS 2023 · 6 citations
- Overcoming Common Flaws in the Evaluation of Selective Classification SystemsJeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner et al.NeurIPS 2024 · 44 citations
- Selective Ensembles for Consistent PredictionsEmily Black, Klas Leino, Matt FredriksonICLR 2022 · 29 citations
- AUC Optimization with a Reject OptionSong-Qing Shen, Bin-Bin Yang, Wei GaoAAAI 2020 · 7 citations
- Confidence-aware Contrastive Learning for Selective ClassificationYu-Chang Wu, Shen-Huan Lyu, Haopu Shang, Xiangyu Wang et al.ICML 2024 · 9 citations
