Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine Translation
Yunlong Liang, Fandong Meng, Jiaan Wang, Jinan Xu, Yufeng Chen, Jie Zhou
摘要
Incrementally expanding the capability of an existing translation model to solve new domain tasks over time is a fundamental and practical problem, which usually suffers from catastrophic forgetting. Generally, multi-domain learning can be seen as a good solution. However, there are two drawbacks: 1) it requires having the training data for all domains available at the same time, which may be unrealistic due to storage or privacy concerns; 2) it requires re-training the model on the data of all domains from scratch when adding a new domain and this is time-consuming and computationally expensive. To address these issues, we present a semi-supervised contrastive distillation framework for incremental neural machine translation. Specifically, to avoid catastrophic forgetting, we propose to exploit unlabeled data from the same distributions of the older domains through knowledge distillation. Further, to ensure the distinct domain characteristics in the model as the number of domains increases, we devise a cross-domain contrastive objective to enhance the distilled knowledge. Extensive experiments on domain translation benchmarks show that our approach, without accessing any previous training data or re-training on all domains from scratch, can significantly prevent the model from forgetting previously learned knowledge while obtaining good performance on the incrementally added domains. 1 Note that the target-side data is not used. 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data ReplayRuiheng Liu, Jinyu Zhang, Yanqi Song, Yu Zhang 等AAAI 2025 · 被引用 5 次
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine TranslationYunlong Liang, Fandong Meng, Jie ZhouACL 2025 · 被引用 1 次
它引用的顶会 Paper14
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Overcoming Catastrophic Forgetting With Unlabeled Data in the WildKibok Lee, Kimin Lee, Jinwoo Shin, Honglak LeeICCV 2019 · 被引用 231 次
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
- Boosting Neural Machine Translation with Similar TranslationsJitao Xu, Josep Maria Crego, Jean SenellartACL 2020 · 被引用 59 次
- Contrastive Distillation on Intermediate Representations for Language Model CompressionSiqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng 等EMNLP 2020 · 被引用 59 次
相关 Paper
- Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine TranslationChenze Shao, Yang FengACL 2022 · 被引用 40 次
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- Asymmetric Synthetic Data Update for Domain Incremental Dataset DistillationMinyoung Oh, Jae-Young SimICLR 2026
- Go From the General to the Particular: Multi-Domain Translation with Domain Transformation NetworksYong Wang, Longyue Wang, Shuming Shi, Victor O. K. Li 等AAAI 2020 · 被引用 30 次
- Incrementer: Transformer for Class-Incremental Semantic Segmentation with Knowledge Distillation Focusing on Old ClassChao Shang, Hongliang Li, Fanman Meng, Qingbo Wu 等CVPR 2023
