On Losses for Modern Language Models
Stephane Aroca-Ouellette, Frank Rudzicz
摘要
BERT set many state-of-the-art results over varied NLU benchmarks by pre-training over two tasks: masked language modelling (MLM) and next sentence prediction (NSP), the latter of which has been highly criticized. In this paper, we 1) clarify NSP's effect on BERT pre-training, 2) explore fourteen possible auxiliary pre-training tasks, of which seven are novel to modern language models, and 3) investigate different ways to include multiple tasks into pre-training. We show that NSP is detrimental to training due to its context splitting and shallow semantic signal. We also identify six auxiliary pre-training tasks -sentence ordering, adjacent sentence prediction, TF prediction, TF-IDF prediction, a Fast-Sent variant, and a Quick Thoughts variant -that outperform a pure MLM baseline. Finally, we demonstrate that using multiple tasks in a multi-task pre-training framework provides better results than using any single auxiliary task. Using these methods, we outperform BERT Base on the GLUE benchmark using fewer than a quarter of the training tokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Should We Be Pre-training? An Argument for End-task Aware Training as an AlternativeLucio M. Dery, Paul Michel, Ameet Talwalkar, Graham NeubigICLR 2022 · 被引用 39 次
- RankGen: Improving Text Generation with Large Ranking ModelsKalpesh Krishna, Yapei Chang, John Wieting, Mohit IyyerEMNLP 2022 · 被引用 27 次
- Multi-CLS BERT: An Efficient Alternative to Traditional EnsemblingHaw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallumACL 2023 · 被引用 8 次
- Multitask Pretraining with Structured Knowledge for Text-to-SQL GenerationRobert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li 等ACL 2023 · 被引用 6 次
- Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task GeneralizationYuxian Gu, Pei Ke, Xiaoyan Zhu, Minlie HuangEMNLP 2022 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto 等ACL 2020 · 被引用 9 次
- GradTS: A Gradient-Based Automatic Auxiliary Task Selection Method Based on Transformer NetworksWeicheng Ma, Renze Lou, Kai Zhang, Lili Wang 等EMNLP 2021 · 被引用 4 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 被引用 105 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
