Span Selection Pre-training for Question Answering
Michael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, G. P. Shrivatsa Bhargav, Dinesh Garg, Avirup Sil
摘要
BERT (Bidirectional Encoder Representations from Transformers) and related pre-trained Transformers have provided large gains across many language understanding tasks, achieving a new state-of-the-art (SOTA). BERT is pretrained on two auxiliary tasks: Masked Language Model and Next Sentence Prediction. In this paper we introduce a new pre-training task inspired by reading comprehension to better align the pre-training from memorization to understanding. Span Selection Pre-Training (SSPT) poses cloze-like training instances, but rather than draw the answer from the model's parameters, it is selected from a relevant passage. We find significant and consistent improvements over both BERT BASE and BERT LARGE on multiple Machine Reading Comprehension (MRC) datasets. Specifically, our proposed model has strong empirical evidence as it obtains SOTA results on Natural Questions, a new benchmark MRC dataset, outperforming BERT LARGE by 3 F1 points on short answer prediction. We also show significant impact in HotpotQA, improving answer prediction F1 by 4 points and supporting fact prediction F1 by 1 point and outperforming the previous best system. Moreover, we show that our pre-training approach is particularly effective when training data is limited, improving the learning curve by a large amount.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah 等EMNLP 2022 · 被引用 63 次
- RikiNet: Reading Wikipedia Pages for Natural Question AnsweringDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan 等ACL 2020 · 被引用 55 次
- Translucent Answer Predictions in Multi-Hop Reading ComprehensionG. P. Shrivatsa Bhargav, Michael R. Glass, Dinesh Garg, Shirish K. Shevade 等AAAI 2020 · 被引用 13 次
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model AdaptationMinki Kang, Moonsu Han, Sung Ju HwangEMNLP 2020 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading ComprehensionLin Ai, Zheng Hui, Zizhou Liu, Julia HirschbergEMNLP 2024 · 被引用 4 次
- On Losses for Modern Language ModelsStephane Aroca-Ouellette, Frank RudziczEMNLP 2020 · 被引用 2 次
- Context-Aware Answer Extraction in Question AnsweringYeon Seonwoo, Ji-Hoon Kim, Jung-Woo Ha, Alice OhEMNLP 2020 · 被引用 32 次
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 被引用 463 次
