Selective Weak Supervision for Neural Information Retrieval
Kaitao Zhang, Chenyan Xiong, Zhenghao Liu, Zhiyuan Liu
摘要
This paper democratizes neural information retrieval to scenarios where large scale relevance training signals are not available. We revisit the classic IR intuition that anchor-document relations approximate query-document relevance and propose a reinforcement weak supervision selection method, ReInfoSelect, which learns to select anchor-document pairs that best weakly supervise the neural ranker (action), using the ranking performance on a handful of relevance labels as the reward. Iteratively, for a batch of anchor-document pairs, ReInfoSelect back propagates the gradients through the neural ranker, gathers its NDCG reward, and optimizes the data selection network using policy gradients, until the neural ranker’s performance peaks on target relevance metrics (convergence). In our experiments on three TREC benchmarks, neural rankers trained by ReInfoSelect, with only publicly available anchor data, significantly outperform feature-based learning to rank methods and match the effectiveness of neural rankers trained with private commercial search logs. Our analyses show that ReInfoSelect effectively selects weak supervision signals based on the stage of the neural ranker training, and intuitively picks anchor-document pairs similar to query-document pairs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning to Augment for Casual User RecommendationJianling Wang, Ya Le, Bo Chang, Yuyan Wang 等WWW 2022 · 被引用 23 次
- MultiCQA: Zero-Shot Transfer of Self-Supervised Text Matching Models on a Massive ScaleAndreas Rücklé, Jonas Pfeiffer, Iryna GurevychEMNLP 2020 · 被引用 2 次
- Leveraging Capsule Routing to Associate Knowledge with Medical Literature HierarchicallyXin Liu, Qingcai Chen, Junying Chen, Wenxiu Zhou 等EMNLP 2021 · 被引用 1 次
- MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module PluginTianshuo Zhou, Sen Mei, Xinze Li, Zhenghao Liu 等ACL 2024
- Few-Shot Text Ranking with Meta Adapted Synthetic Weak SupervisionSi Sun, Yingzhuo Qian, Zhenghao Liu, Chenyan Xiong 等ACL 2021
相关 Paper
- Reward-based Input Construction for Cross-document Relation ExtractionByeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul MoonACL 2024 · 被引用 3 次
- Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-RankingTim Baumgärtner, Leonardo F. R. Ribeiro, Nils Reimers, Iryna GurevychEMNLP 2022 · 被引用 3 次
- PSLOG: Pretraining with Search Logs for Document RankingZhan Su, Zhicheng Dou, Yujia Zhou, Ziyuan Zhao 等KDD 2023 · 被引用 2 次
- Enhancing Generative Retrieval with Reinforcement Learning from Relevance FeedbackYujia Zhou, Zhicheng Dou, Ji-Rong WenEMNLP 2023 · 被引用 14 次
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
