Churn Reduction via Distillation
Heinrich Jiang, Harikrishna Narasimhan, Dara Bahri, Andrew Cotter, Afshin Rostamizadeh
摘要
In real-world systems, models are frequently updated as more data becomes available, and in addition to achieving high accuracy, the goal is to also maintain a low difference in predictions compared to the base model (i.e. predictive “churn”). If model retraining results in vastly different behavior, then it could cause negative effects in downstream systems, especially if this churn can be avoided with limited impact on model accuracy. In this paper, we show an equivalence between training with distillation using the base model as the teacher and training with an explicit constraint on the predictive churn. We then show that distillation performs strongly for low churn training against a number of recent baselines on a wide range of datasets and model architectures, including fully-connected networks, convolutional networks, and transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Assessing Generalization of SGD via DisagreementYiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico KolterICLR 2022 · 被引用 134 次
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
它引用的顶会 Paper3
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 被引用 288 次
相关 Paper
- Locally Adaptive Label Smoothing Improves Predictive ChurnDara Bahri, Heinrich JiangICML 2021 · 被引用 16 次
- Positive-Congruent Training: Towards Regression-Free Model UpdatesSijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang 等CVPR 2021
- Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher KnowledgeFreya Behrens, Lenka ZdeborováICLR 2026 · 被引用 9 次
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su 等NeurIPS 2022 · 被引用 14 次
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram 等ICML 2025
