Measuring and Reducing Model Update Regression in Structured Prediction for NLP
Deng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su, Lei Shu, Yi Zhang
Abstract
Recent advance in deep learning has led to the rapid adoption of machine learning-based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward compatibility requires that the new model does not regress on cases that were correctly handled by its predecessor. This work studies model update regression in structured prediction tasks. We choose syntactic dependency parsing and conversational semantic parsing as representative examples of structured prediction tasks in NLP. First, we measure and analyze model update regression in different model update settings. Next, we explore and benchmark existing techniques for reducing model update regression including model ensemble and knowledge distillation. We further propose a simple and effective method, Backward-Congruent Re-ranking (BCR), by taking into account the characteristics of structured prediction. Experiments show that BCR can better mitigate model update regression than model ensemble and knowledge distillation approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 907036f6-58e5-4d53-bf42-3e72433f0cc0Cited by top-tier papers3
- FlipGuard: Defending Preference Alignment against Update Regression with Constrained OptimizationMingye Zhu, Yi Liu, Quan Wang, Junbo Guo et al.EMNLP 2024 · 2 citations
- Mitigating Negative Flips via Margin Preserving TrainingSimone Ricci, Niccolò Biondi, Federico Pernici, Alberto Del BimboAAAI 2026
- Pseudo-label Training and Model Inertia in Neural Machine TranslationBenjamin Hsu, Anna Currey, Xing Niu, Maria Nadejde et al.ICLR 2023
Builds on9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Backward-Compatible Prediction Updates: A Probabilistic ApproachFrederik Träuble, Julius von Kügelgen, Matthäus Kleindessner, Francesco Locatello et al.NeurIPS 2021 · 20 citations
- Neural Reranking for Dependency Parsing: An EvaluationBich-Ngoc Do, Ines RehbeinACL 2020 · 9 citations
- On Finding the K-best Non-projective Dependency TreesRan Zmigrod, Tim Vieira, Ryan CotterellACL 2021
Related papers
- Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model UpdatesYuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang et al.ACL 2021
- Adversarial Attack and Defense of Structured Prediction ModelsWenjuan Han, Liwen Zhang, Yong Jiang, Kewei TuEMNLP 2020 · 32 citations
- Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency ParsingZhuoran Li, Chunming Hu, Junfan Chen, Zhijun Chen et al.AAAI 2025 · 1 citation
- A Generalized Backward Compatibility MetricTomoya SakaiKDD 2022 · 3 citations
- Towards Backward-Compatible Representation LearningYantao Shen, Yuanjun Xiong, Wei Xia, Stefano SoattoCVPR 2020
