When is Multicalibration Post-Processing Necessary?
Dutch Hansen, Siddartha Devic, Preetum Nakkiran, Vatsal Sharan
Abstract
Calibration is a well-studied property of predictors which guarantees meaningful uncertainty estimates. Multicalibration is a related notion -- originating in algorithmic fairness -- which requires predictors to be simultaneously calibrated over a potentially complex and overlapping collection of protected subpopulations (such as groups defined by ethnicity, race, or income). We conduct the first comprehensive study evaluating the usefulness of multicalibration post-processing across a broad set of tabular, image, and language datasets for models spanning from simple decision trees to 90 million parameter fine-tuned LLMs. Our findings can be summarized as follows: (1) models which are calibrated out of the box tend to be relatively multicalibrated without any additional post-processing; (2) multicalibration post-processing can help inherently uncalibrated models and large vision and language models; and (3) traditional calibration measures may sometimes provide multicalibration implicitly. More generally, we also distill many independent observations which may be useful for practical and effective applications of multicalibration post-processing in real-world contexts. We also release a python package implementing multicalibration algorithms, available via `pip install multicalibration'.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a29fe820-2e13-4a22-92e9-bf69ca99ca19Cited by top-tier papers8
- Discretization-free Multicalibration through Loss Minimization over Tree EnsemblesHongyi Henry Jin, Zijun Ding, Dung Daniel T. Ngo, Zhiwei Steven WuNeurIPS 2025 · 6 citations
- Selective Omniprediction and Fair AbstentionSílvia Casacuberta, Varun KanadeNeurIPS 2025 · 3 citations
- On Group Sufficiency Under Label BiasHaoran Zhang, Olawale Salaudeen, Marzyeh GhassemiNeurIPS 2025 · 2 citations
- Who's the (Multi-)Fairest of Them All: Rethinking Interpolation-Based Data Augmentation Through the Lens of MulticalibrationKarina Halevy, Karly Hou, Charumathi BadrinathAAAI 2025 · 2 citations
- How Global Calibration Strengthens MultiaccuracySílvia Casacuberta, Parikshit Gopalan, Varun Kanade, Omer ReingoldFOCS 2025 · 1 citation
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Calibrating Predictions to Decisions: A Novel Approach to Multi-Class CalibrationShengjia Zhao, Michael P. Kim, Roshni Sahoo, Tengyu Ma et al.NeurIPS 2021 · 96 citations
Related papers
- Fair Risk Control: A Generalized Framework for Calibrating Multi-group Fairness RisksLujing Zhang, Aaron Roth, Linjun ZhangICML 2024 · 11 citations
- Omnipredictors for Constrained OptimizationLunjia Hu, Inbal Rachel Livni Navon, Omer Reingold, Chutong YangICML 2023 · 17 citations
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 39 citations
- Sample Complexity of Uniform Convergence for MulticalibrationEliran Shabat, Lee Cohen, Yishay MansourNeurIPS 2020 · 32 citations
- FairCal: Fairness Calibration for Face VerificationTiago Salvador, Stephanie Cairns, Vikram Voleti, Noah Marshall et al.ICLR 2022 · 21 citations
