Churn Reduction via Distillation
Heinrich Jiang, Harikrishna Narasimhan, Dara Bahri, Andrew Cotter, Afshin Rostamizadeh
Abstract
In real-world systems, models are frequently updated as more data becomes available, and in addition to achieving high accuracy, the goal is to also maintain a low difference in predictions compared to the base model (i.e. predictive “churn”). If model retraining results in vastly different behavior, then it could cause negative effects in downstream systems, especially if this churn can be avoided with limited impact on model accuracy. In this paper, we show an equivalence between training with distillation using the base model as the teacher and training with an explicit constraint on the predictive churn. We then show that distillation performs strongly for low churn training against a number of recent baselines on a wide range of datasets and model architectures, including fully-connected networks, convolutional networks, and transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 134 citations
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
Builds on3
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
Related papers
- Locally Adaptive Label Smoothing Improves Predictive ChurnDara Bahri, Heinrich JiangICML 2021 · 16 citations
- Positive-Congruent Training: Towards Regression-Free Model UpdatesSijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang et al.CVPR 2021
- Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher KnowledgeFreya Behrens, Lenka ZdeborováICLR 2026 · 9 citations
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su et al.NeurIPS 2022 · 14 citations
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram et al.ICML 2025
