Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering
Jiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen
Abstract
To alleviate the data scarcity problem in training question answering systems, recent works propose additional intermediate pre-training for dense passage retrieval (DPR). However, there still remains a large discrepancy between the provided upstream signals and the downstream question-passage relevance, which leads to less improvement. To bridge this gap, we propose the HyperLink-induced Pre-training (HLP), a method to pre-train the dense retriever with the text relevance induced by hyperlink-based topology within Web documents. We demonstrate that the hyperlink-based structures of dual-link and co-mention can provide effective relevance signals for large-scale pre-training that better facilitate downstream passage retrieval. We investigate the effectiveness of our approach across a wide range of open-domain QA datasets under zero-shot, few-shot, multi-hop, and out-of-domain scenarios. The experiments show our HLP outperforms the BM25 by up to 7 points as well as other pre-training methods by more than 10 points in terms of top-20 retrieval accuracy under the zero-shot scenario. Furthermore, HLP significantly outperforms other pre-training methods under the other scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c105297f-bf90-4809-9250-ba0afdb5e24dCited by top-tier papers8
- LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale RetrievalTao Shen, Xiubo Geng, Chongyang Tao, Can Xu et al.ICLR 2023 · 14 citations
- Weaker Than You Think: A Critical Look at Weakly Supervised LearningDawei Zhu, Xiaoyu Shen, Marius Mosbach, Andreas Stephan et al.ACL 2023 · 13 citations
- Chain-of-Skills: A Configurable Model for Open-Domain Question AnsweringKaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu et al.ACL 2023 · 12 citations
- Synergistic Interplay between Search and Large Language Models for Information RetrievalJiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen et al.ACL 2024 · 7 citations
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong et al.EMNLP 2022 · 6 citations
Builds on5
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- End-to-End Training of Neural Retrievers for Open-Domain Question AnsweringDevendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant et al.ACL 2021
Related papers
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
- Improving Passage Retrieval with Zero-Shot Question GenerationDevendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan et al.EMNLP 2022 · 69 citations
- HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval GeneralizationZefeng Cai, Chongyang Tao, Tao Shen, Can Xu et al.ICLR 2023
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- Few-shot Reranking for Multi-hop QA via Language Model PromptingMuhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee et al.ACL 2023 · 5 citations
