Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates
Yuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang, Stefano Soatto
Abstract
Behavior of deep neural networks can be inconsistent between different versions. Regressions 1 during model update are a common cause of concern that often over-weigh the benefits in accuracy or efficiency gain. This work focuses on quantifying, reducing and analyzing regression errors in the NLP model updates. Using negative flip rate as regression measure, we show that regression has a prevalent presence across tasks in the GLUE benchmark. We formulate the regression-free model updates into a constrained optimization problem, and further reduce it into a relaxed form which can be approximately optimized through knowledge distillation training method. We empirically analyze how model ensemble reduces regression. Finally, we conduct CHECKLIST behavioral testing to understand the distribution of regressions across linguistic phenomena, and the efficacy of ensemble and distillation methods. * * Work done while at Amazon AWS AI. 1 Here regression refers to bugs in software testing instead of the statistical estimation method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7af89491-f239-4627-a331-98e2960fa270Cited by top-tier papers6
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su et al.NeurIPS 2022 · 14 citations
- Lightweight Approaches to DNN Regression Error Reduction: An Uncertainty Alignment PerspectiveZenan Li, Maorun Zhang, Jingwei Xu, Yuan Yao et al.ICSE 2023 · 3 citations
- FlipGuard: Defending Preference Alignment against Update Regression with Constrained OptimizationMingye Zhu, Yi Liu, Quan Wang, Junbo Guo et al.EMNLP 2024 · 2 citations
- RegTrieve: Reducing System-Level Regression Errors for Machine Learning Systems via Retrieval-Enhanced EnsembleJunming Cao, Xuwen Xiang, Mingfei Cheng, Bihuan Chen et al.FSE 2025
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
Builds on8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Positive-Congruent Training: Towards Regression-Free Model UpdatesSijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang et al.CVPR 2021
- It Takes Two to EntangleZhanghan Wang, Ding Ding, Hang Zhu, Haibin Lin et al.ASPLOS 2026
- Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmShaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang et al.ACL 2022
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu et al.ICSE 2026
