Counterfactually Comparing Abstaining Classifiers
Yo Joong Choe, Aditya Gangrade, Aaditya Ramdas
Abstract
Abstaining classifiers have the option to abstain from making predictions on inputs that they are unsure about. These classifiers are becoming increasingly popular in high-stake decision-making problems, as they can withhold uncertain predictions to improve their reliability and safety. When evaluating black-box abstaining classifier(s), however, we lack a principled approach that accounts for what the classifier would have predicted on its abstentions. These missing predictions are crucial when, e.g., a radiologist is unsure of their diagnosis or when a driver is inattentive in a self-driving car. In this paper, we introduce a novel approach and perspective to the problem of evaluating and comparing abstaining classifiers by treating abstentions as missing data. Our evaluation approach is centered around defining the counterfactual score of an abstaining classifier, defined as the expected performance of the classifier had it not been allowed to abstain. We specify the conditions under which the counterfactual score is identifiable: if the abstentions are stochastic, and if the evaluation data is independent of the training data (ensuring that the predictions are missing at random), then the score is identifiable. Note that, if abstentions are deterministic, then the score is unidentifiable because the classifier can perform arbitrarily poorly on its abstentions. Leveraging tools from observational causal inference, we then develop nonparametric and doubly robust methods to efficiently estimate this quantity under identification. Our approach is examined in both simulated and real data experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87870b60-b54e-4fc3-bad1-4562fc7e1740Builds on1
Related papers
- LLMs (Almost) Never Abstain Under Medical UncertaintyAlessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Gianluca MoroACL 2026
- Training Private Models That Know What They Don't KnowStephan Rabanser, Anvith Thudi, Abhradeep Guha Thakurta, Krishnamurthy Dvijotham et al.NeurIPS 2023 · 10 citations
- Treatment Responder Classification with AbstentionHaoxiang Wang, Haoxuan Li, Ziyan Wang, Zhiheng Zhang et al.ICML 2026
- Bounded-Abstention Pairwise Learning to RankAntonio Ferrara, Andrea Pugnana, Francesco Bonchi, Salvatore RuggieriKDD 2026
- AUC Optimization with a Reject OptionSong-Qing Shen, Bin-Bin Yang, Wei GaoAAAI 2020 · 7 citations
