Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting
Wangchunshu Zhou, Tao Ge, Canwen Xu, Ke Xu, Furu Wei
Abstract
In this paper, we propose Sequence Span Rewriting (SSR), a self-supervised task for sequence-to-sequence (Seq2Seq) pre-training. SSR learns to refine the machine-generated imperfect text spans into ground truth text. SSR provides more fine-grained and informative supervision in addition to the original textinfilling objective. Compared to the prevalent text infilling objectives for Seq2Seq pretraining, SSR is naturally more consistent with many downstream generation tasks that require sentence rewriting (e.g., text summarization, question generation, grammatical error correction, and paraphrase generation). We conduct extensive experiments by using SSR to improve the typical Seq2Seq pre-trained model T5 in a continual pre-training setting and show substantial improvements over T5 on various natural language generation tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ff69d03-c7b6-4f39-b94d-1c9ee9b94b91Cited by top-tier papers4
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 114 citations
- VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingWangchunshu Zhou, Yan Zeng, Shizhe Diao, Xinsong ZhangICML 2022 · 17 citations
- EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq GenerationTao Ge, Si-Qing Chen, Furu WeiEMNLP 2022 · 16 citations
- Write and Paint: Generative Vision-Language Models are Unified Modal LearnersShizhe Diao, Wangchunshu Zhou, Xinsong Zhang, Jiawei WangICLR 2023 · 5 citations
Builds on10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Self-Supervised Query Reformulation for Code SearchYuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong GuFSE 2023 · 14 citations
- Copy That! Editing Sequences by Copying SpansSheena Panthaplackel, Miltiadis Allamanis, Marc BrockschmidtAAAI 2021 · 28 citations
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge et al.ICSE 2022 · 99 citations
- mT6: Multilingual Pretrained Text-to-Text Transformer with Translation PairsZewen Chi, Li Dong, Shuming Ma, Shaohan Huang et al.EMNLP 2021 · 60 citations
- Generating Sequences by Learning to Self-CorrectSean Welleck, Ximing Lu, Peter West, Faeze Brahman et al.ICLR 2023 · 30 citations
