RINK: Reader-Inherited Evidence Reranker for Table-and-Text Open Domain Question Answering
Eunhwan Park, Sung-Min Lee, Daeryong Seo, Seonhoon Kim, Inho Kang, Seung-Hoon Na
Abstract
Most approaches used in open-domain question answering on hybrid data that comprises both tabular-and-textual contents are based on a Retrieval-Reader pipeline in which the retrieval module finds relevant "heterogenous" evidence for a given question and the reader module generates an answer from the retrieved evidence. In this paper, we present a Retriever-Reranker-Reader framework by newly proposing a Reader-INherited evidence reranKer (RINK) where a reranker module is designed by finetuning the reader's neural architecture based on a simple prompting method. Our underlying assumption of reusing the reader's module for the reranker is that the reader's ability to generating an answer from evidence contains the knowledge required for the reranking, because the reranker needs to "read" in-depth a question and evidences more carefully and elaborately than a baseline retriever. Furthermore, we present a simple and effective pretraining method by extensively deploying the commonly used data augmentation methods of cell corruption and cell reordering based on the pretraining taskstabular-and-textual entailment and cross-modal masked language modeling. Experimental results on OTT-QA, a largescale table -and-text open-domain question answering dataset, show that the proposed RINK armed with our pretraining procedure makes improvements over the baseline reranking method and leads to state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cd0154f-24e3-41d0-a0f9-0d88593be28dBuilds on14
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
Related papers
- Open Question Answering over Tables and TextWenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang et al.ICLR 2021 · 76 citations
- UnitedQA: A Hybrid Approach for Open Domain Question AnsweringHao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He et al.ACL 2021
- You Only Need One Model for Open-domain Question AnsweringHaejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape et al.EMNLP 2022
- Open Domain Question Answering with A Unified Knowledge InterfaceKaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg et al.ACL 2022 · 45 citations
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
