Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates
Yuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang, Stefano Soatto
摘要
Behavior of deep neural networks can be inconsistent between different versions. Regressions 1 during model update are a common cause of concern that often over-weigh the benefits in accuracy or efficiency gain. This work focuses on quantifying, reducing and analyzing regression errors in the NLP model updates. Using negative flip rate as regression measure, we show that regression has a prevalent presence across tasks in the GLUE benchmark. We formulate the regression-free model updates into a constrained optimization problem, and further reduce it into a relaxed form which can be approximately optimized through knowledge distillation training method. We empirically analyze how model ensemble reduces regression. Finally, we conduct CHECKLIST behavioral testing to understand the distribution of regressions across linguistic phenomena, and the efficacy of ensemble and distillation methods. * * Work done while at Amazon AWS AI. 1 Here regression refers to bugs in software testing instead of the statistical estimation method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su 等NeurIPS 2022 · 被引用 14 次
- Lightweight Approaches to DNN Regression Error Reduction: An Uncertainty Alignment PerspectiveZenan Li, Maorun Zhang, Jingwei Xu, Yuan Yao 等ICSE 2023 · 被引用 3 次
- FlipGuard: Defending Preference Alignment against Update Regression with Constrained OptimizationMingye Zhu, Yi Liu, Quan Wang, Junbo Guo 等EMNLP 2024 · 被引用 2 次
- RegTrieve: Reducing System-Level Regression Errors for Machine Learning Systems via Retrieval-Enhanced EnsembleJunming Cao, Xuwen Xiang, Mingfei Cheng, Bihuan Chen 等FSE 2025
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
它引用的顶会 Paper8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- Positive-Congruent Training: Towards Regression-Free Model UpdatesSijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang 等CVPR 2021
- It Takes Two to EntangleZhanghan Wang, Ding Ding, Hang Zhu, Haibin Lin 等ASPLOS 2026
- Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmShaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang 等ACL 2022
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 等ICLR 2021 · 被引用 90 次
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu 等ICSE 2026
