BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense Retrieval
Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng
摘要
Dense retrieval has shown promise in the first-stage retrieval process when trained on in-domain labeled datasets. However, previous studies have found that dense retrieval is hard to generalize to unseen domains due to its weak modeling of domain-invariant and interpretable feature (i.e., matching signal between two texts, which is the essence of information retrieval). In this paper, we propose a novel method to improve the generalization of dense retrieval via capturing matching signal called BERM. Fully fine-grained expression and query-oriented saliency are two properties of the matching signal. Thus, in BERM, a single passage is segmented into multiple units and two unit-level requirements are proposed for representation as the constraint in training to obtain the effective matching signal. One is semantic unit balance and the other is essential matching unit extractability. Unit-level view and balanced semantics make representation express the text in a fine-grained manner. Essential matching unit extractability makes passage representation sensitive to the given query to extract the pure matching information from the passage containing complex context. Experiments on BEIR show that our method can be effectively combined with different dense retrieval training methods (vanilla, hard negatives mining and knowledge distillation) to improve its generalization ability without any additional inference overhead and target domain data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive TasksShicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng 等WWW 2024 · 被引用 104 次
- Neural Retrievers are Biased Towards LLM-Generated ContentSunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu 等KDD 2024 · 被引用 26 次
- MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation SystemJihao Zhao, Zhiyuan Ji, Zhaoxin Fan, Hanyu Wang 等ACL 2025 · 被引用 21 次
- List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented GenerationShicheng Xu, Liang Pang, Jun Xu, Huawei Shen 等WWW 2024 · 被引用 13 次
- Constrained Auto-Regressive Decoding Constrains Generative RetrievalShiguang Wu, Zhaochun Ren, Xin Xin, Jiyuan Yang 等SIGIR 2025 · 被引用 3 次
它引用的顶会 Paper9
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
- Towards a Theoretical Framework of Out-of-Distribution GeneralizationHaotian Ye, Chuanlong Xie, Tianle Cai, Ruichen Li 等NeurIPS 2021 · 被引用 159 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
相关 Paper
- LED: Lexicon-Enlightened Dense Retriever for Large-Scale RetrievalKai Zhang, Chongyang Tao, Tao Shen, Can Xu 等WWW 2023 · 被引用 27 次
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 被引用 63 次
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document RetrievalWonbin Kweon, Runchu Tian, Seongku Kang, Pengcheng Jiang 等WWW 2026
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang 等EMNLP 2022 · 被引用 21 次
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen 等ICLR 2026 · 被引用 3 次
