Transfer Learning for Sequence Generation: from Single-source to Multi-source
Xuancheng Huang, Jingfang Xu, Maosong Sun, Yang Liu
Abstract
Multi-source sequence generation (MSG) is an important kind of sequence generation tasks that takes multiple sources, including automatic post-editing, multi-source translation, multi-document summarization, etc. As MSG tasks suffer from the data scarcity problem and recent pretrained models have been proven to be effective for low-resource downstream tasks, transferring pretrained sequenceto-sequence models to MSG tasks is essential. Although directly finetuning pretrained models on MSG tasks and concatenating multiple sources into a single long sequence is regarded as a simple method to transfer pretrained models to MSG tasks, we conjecture that the direct finetuning method leads to catastrophic forgetting and solely relying on pretrained selfattention layers to capture cross-source information is not sufficient. Therefore, we propose a two-stage finetuning method to alleviate the pretrain-finetune discrepancy and introduce a novel MSG model with a fine encoder to learn better representations in MSG tasks. Experiments show that our approach achieves new state-of-the-art results on the WMT17 APE task and multi-source translation task using the WMT14 test set. When adapted to documentlevel translation, our framework outperforms strong baselines significantly. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3bbcc94-5fa0-4525-98af-d0877825b432Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut et al.ACL 2020 · 168 citations
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni et al.EMNLP 2020 · 104 citations
- Incorporating BERT into Parallel Sequence Decoding with AdaptersJunliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei et al.NeurIPS 2020 · 72 citations
Related papers
- Cross-Lingual Natural Language Generation via Pre-TrainingZewen Chi, Li Dong, Furu Wei, Wenhui Wang et al.AAAI 2020 · 142 citations
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan et al.NeurIPS 2020 · 165 citations
- MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better TranslatorsZhixing Tan, Xiangwen Zhang, Shuo Wang, Yang LiuACL 2022 · 58 citations
- Transitional Adaptation of Pretrained Models for Visual StorytellingYoungjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim et al.CVPR 2021
- Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionFei Yuan, Linjun Shou, Xuanyu Bai, Ming Gong et al.ACL 2020 · 21 citations
