Transfer Learning for Sequence Generation: from Single-source to Multi-source
Xuancheng Huang, Jingfang Xu, Maosong Sun, Yang Liu
摘要
Multi-source sequence generation (MSG) is an important kind of sequence generation tasks that takes multiple sources, including automatic post-editing, multi-source translation, multi-document summarization, etc. As MSG tasks suffer from the data scarcity problem and recent pretrained models have been proven to be effective for low-resource downstream tasks, transferring pretrained sequenceto-sequence models to MSG tasks is essential. Although directly finetuning pretrained models on MSG tasks and concatenating multiple sources into a single long sequence is regarded as a simple method to transfer pretrained models to MSG tasks, we conjecture that the direct finetuning method leads to catastrophic forgetting and solely relying on pretrained selfattention layers to capture cross-source information is not sufficient. Therefore, we propose a two-stage finetuning method to alleviate the pretrain-finetune discrepancy and introduce a novel MSG model with a fine encoder to learn better representations in MSG tasks. Experiments show that our approach achieves new state-of-the-art results on the WMT17 APE task and multi-source translation task using the WMT14 test set. When adapted to documentlevel translation, our framework outperforms strong baselines significantly. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He 等ICLR 2020 · 被引用 391 次
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni 等EMNLP 2020 · 被引用 104 次
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei 等NeurIPS 2020 · 被引用 72 次
相关 Paper
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang 等AAAI 2020 · 被引用 142 次
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan 等NeurIPS 2020 · 被引用 165 次
- MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better TranslatorsZhixing Tan, Xiangwen Zhang, Shuo Wang, Yang LiuACL 2022 · 被引用 58 次
- Transitional Adaptation of Pretrained Models for Visual StorytellingYoungjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim 等CVPR 2021
- Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionFei Yuan, Linjun Shou, Xuanyu Bai, Ming Gong 等ACL 2020 · 被引用 21 次
