Span Selection Pre-training for Question Answering
Michael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, G. P. Shrivatsa Bhargav, Dinesh Garg, Avirup Sil
Abstract
BERT (Bidirectional Encoder Representations from Transformers) and related pre-trained Transformers have provided large gains across many language understanding tasks, achieving a new state-of-the-art (SOTA). BERT is pretrained on two auxiliary tasks: Masked Language Model and Next Sentence Prediction. In this paper we introduce a new pre-training task inspired by reading comprehension to better align the pre-training from memorization to understanding. Span Selection Pre-Training (SSPT) poses cloze-like training instances, but rather than draw the answer from the model's parameters, it is selected from a relevant passage. We find significant and consistent improvements over both BERT BASE and BERT LARGE on multiple Machine Reading Comprehension (MRC) datasets. Specifically, our proposed model has strong empirical evidence as it obtains SOTA results on Natural Questions, a new benchmark MRC dataset, outperforming BERT LARGE by 3 F1 points on short answer prediction. We also show significant impact in HotpotQA, improving answer prediction F1 by 4 points and supporting fact prediction F1 by 1 point and outperforming the previous best system. Moreover, we show that our pre-training approach is particularly effective when training data is limited, improving the learning curve by a large amount.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20dea636-3aa7-42aa-8e2e-3e708691fb70Cited by top-tier papers10
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial DomainRaj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah et al.EMNLP 2022 · 63 citations
- RikiNet: Reading Wikipedia Pages for Natural Question AnsweringDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan et al.ACL 2020 · 55 citations
- Translucent Answer Predictions in Multi-Hop Reading ComprehensionG. P. Shrivatsa Bhargav, Michael R. Glass, Dinesh Garg, Shirish K. Shevade et al.AAAI 2020 · 13 citations
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model AdaptationMinki Kang, Moonsu Han, Sung Ju HwangEMNLP 2020 · 12 citations
Builds on1
Related papers
- Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading ComprehensionLin Ai, Zheng Hui, Zizhou Liu, Julia HirschbergEMNLP 2024 · 4 citations
- On Losses for Modern Language ModelsStephane Aroca-Ouellette, Frank RudziczEMNLP 2020 · 2 citations
- Context-Aware Answer Extraction in Question AnsweringYeon Seonwoo, Ji-Hoon Kim, Jung-Woo Ha, Alice OhEMNLP 2020 · 32 citations
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen et al.ACL 2020 · 39 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
