QuASE: Question-Answer Driven Sentence Encoding
Hangfeng He, Qiang Ning, Dan Roth
Abstract
Question-answering (QA) data often encodes essential information in many facets. This paper studies a natural question: Can we get supervision from QA data for other tasks (typically, non-QA ones)? For example, can we use QAMR (Michael et al., 2017) to improve named entity recognition? We suggest that simply further pre-training BERT is often not the best option, and propose the question-answer driven sentence encoding (QuASE) framework. QuASE learns representations from QA data, using BERT or other state-of-the-art contextual language models. In particular, we observe the need to distinguish between two types of sentence encodings, depending on whether the target task is a single- or multi-sentence input; in both cases, the resulting encoding is shown to be an easy-to-use plugin for many downstream tasks. This work may point out an alternative way to supervise NLP tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng et al.EMNLP 2020 · 79 citations
- QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and BaselinesValentina Pyatkin, Ayal Klein, Reut Tsarfaty, Ido DaganEMNLP 2020 · 33 citations
- Peek Across: Improving Multi-Document Modeling via Cross-Document Question-AnsweringAvi Caciularu, Matthew E. Peters, Jacob Goldberger, Ido Dagan et al.ACL 2023 · 8 citations
- Foreseeing the Benefits of Incidental SupervisionHangfeng He, Mingyuan Zhang, Qiang Ning, Dan RothEMNLP 2021 · 7 citations
- QASem Parsing: Text-to-text Modeling of QA-based SemanticsAyal Klein, Eran Hirsch, Ron Eliav, Valentina Pyatkin et al.EMNLP 2022 · 4 citations
Related papers
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language ModelWenhan Xiong, Jingfei Du, William Yang Wang, Veselin StoyanovICLR 2020 · 215 citations
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 85 citations
- On Losses for Modern Language ModelsStephane Aroca-Ouellette, Frank RudziczEMNLP 2020 · 2 citations
- Cross-Thought for Sentence Encoder Pre-trainingShuohang Wang, Yuwei Fang, Siqi Sun, Zhe Gan et al.EMNLP 2020 · 17 citations
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto et al.ACL 2020 · 9 citations
