Positive-Congruent Training: Towards Regression-Free Model Updates
Sijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang, Siqi Deng, Meng Wang, Wei Xia, Stefano Soatto
Abstract
Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correctly classified by the old (reference) model. Positivecongruent (PC) training aims at reducing error rate while at the same time reducing negative flips, thus maximizing congruency with the reference model only on positive predictions, unlike model distillation. We propose a simple approach for PC training, Focal Distillation, which enforces congruence with the reference model by giving more weights to samples that were correctly classified. We also found that, if the reference model itself can be chosen as an ensemble of multiple deep neural networks, negative flips can be further reduced without affecting the new model's accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe0efe20-08e3-4252-91f9-e68cfb2563a7Cited by top-tier papers18
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 84 citations
- Backward-Compatible Prediction Updates: A Probabilistic ApproachFrederik Träuble, Julius von Kügelgen, Matthäus Kleindessner, Francesco Locatello et al.NeurIPS 2021 · 20 citations
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su et al.NeurIPS 2022 · 14 citations
- Fixes That Fail: Self-Defeating Improvements in Machine-Learning SystemsRuihan Wu, Chuan Guo, Awni Y. Hannun, Laurens van der MaatenNeurIPS 2021 · 13 citations
- λ-Orthogonality Regularization for Compatible Representation LearningSimone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras et al.NeurIPS 2025 · 8 citations
Builds on1
Related papers
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
- Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model UpdatesYuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang et al.ACL 2021
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 205 citations
