Improving Model Probability Calibration by Integration of Large Data Sources with Biased Labels
Renat Sergazinov, Richard Chen, Cheng Ji, Jing Wu, Daniel Cociorva, Hakan Brunzell
摘要
Probability calibration transforms raw output of a classification model into empirically interpretable probability. When the model is purposed to detect rare event and only a small expensive data source has clean labels, it becomes extraordinarily challenging to obtain accurate probability calibration. Utilizing an additional large cheap data source is very helpful, however, such data sources oftentimes suffer from biased labels. To this end, we introduce an approximate expectation-maximization (EM) algorithm to extract useful information from the large data sources. For a family of calibration methods based on the logistic likelihood, we derive closed-form updates and call the resulting iterative algorithm CalEM. We show that CalEM inherits convergence guarantees from the approximate EM algorithm. We test the proposed model in simulation and on the real marketing datasets, where it shows significant performance increases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- How Well Calibrated are Extreme Multi-label Classifiers? An Empirical AnalysisNasib Ullah, Erik Schultheis, Jinbin Zhang, Rohit BabbarKDD 2025 · 被引用 1 次
- Fill In The Gaps: Model Calibration and Generalization with Synthetic DataYang Ba, Michelle Mancenido, Rong PanEMNLP 2024 · 被引用 2 次
- Obtaining Calibrated Probabilities with Personalized Ranking ModelsWonbin Kweon, SeongKu Kang, Hwanjo YuAAAI 2022 · 被引用 20 次
- Evidential Deep Partial Label Learning to Quantify Disambiguation UncertaintyJinfu Fan, Jiangnan Li, Xiaohui Zhong, Kangrui Ren 等CVPR 2026 · 被引用 3 次
- From Noisy Prediction to True Label: Noisy Prediction Calibration via Generative ModelHeeSun Bae, Seungjae Shin, Byeonghu Na, JoonHo Jang 等ICML 2022 · 被引用 29 次
