Embedding Prior Task-specific Knowledge into Language Models for Context-aware Document Ranking
Shuting Wang, Yutao Zhu, Zhicheng Dou
Abstract
Exploiting users' contextual behaviors in the current session has been proven favorable to the document ranking task. Recently, the context-aware document ranking task has benefited from pre-trained language models (PLMs) due to their superior ability in language modeling. Most PLM-based context-aware document ranking models implicitly learn task-specific knowledge by fine-tuning PLMs on historical search logs. However, since search log data is noisy and contains various user intents and search patterns, such a black-box way may prevent models from fully mastering effective context-aware search knowledge. To solve this problem, we propose LOCK, a PLM-based context-aware document ranking model that explicitly embeds task-specific prior knowledge into PLMs to guide the model optimization. From local to global, we identify three types of task-specific knowledge, including intra-turn signals, inter-turn signals, and global session signals. LOCK formulates such prior knowledge into prior attention biases for impacting the fine-tuning of PLMs. This operation can guide the ranking model by task-specific prior knowledge, thereby improving model convergence and ranking ability. Additionally, we introduce a task-specific pre-training stage that involves masked language modeling and the soft reconstruction of the prior attention matrix, which helps the PLMs adapt to our task. Extensive experiments validate the effectiveness and convergence of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84e9b3c0-c172-46a1-b246-ee90289339faBuilds on3
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 105 citations
- Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching TasksTingyu Xia, Yue Wang, Yuan Tian, Yi ChangWWW 2021 · 56 citations
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 55 citations
Related papers
- PSLOG: Pretraining with Search Logs for Document RankingZhan Su, Zhicheng Dou, Yujia Zhou, Ziyuan Zhao et al.KDD 2023 · 2 citations
- Encoding History with Context-aware Representation Learning for Personalized SearchYujia Zhou, Zhicheng Dou, Ji-Rong WenSIGIR 2020 · 56 citations
- Leveraging Passage-level Cumulative Gain for Document RankingZhijing Wu, Jiaxin Mao, Yiqun Liu, Jingtao Zhan et al.WWW 2020 · 40 citations
- KALM: Knowledge-Aware Integration of Local, Document, and Global Contexts for Long Document UnderstandingShangbin Feng, Zhaoxuan Tan, Wenqian Zhang, Zhenyu Lei et al.ACL 2023 · 5 citations
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou et al.AAAI 2021 · 57 citations
