ElicitR: Unlocking Latent Reasoning in Dense Retrievers via Generative Regularization
Fengyu Cai, Iryna Gurevych, Heinz Koeppl
摘要
Reasoning-intensive retrieval is increasingly important for downstream applications, requiring more than lexical overlap or coarse semantic matching. While prior work mainly relies on Language Models (LMs) to synthesize reasoning-oriented supervision, we posit that it is already latent in LM-based retrievers but suppressed by contrastive overfitting. To elicit this latent reasoning, we introduce ElicitR, a retriever–LM framework with generative regularization that captures nuanced relationships among a query and its candidate documents beyond binary relevance. Concretely, alongside contrastive learning, we regularize the retriever by co-training a small LM on query–positive–negative batches. Next token prediction (NTP) for each text is conditioned on its prefix and the other in-batch texts, with cross-text conditioning weighted by retriever-computed similarities. Using MS MARCO as the only paired query-document supervision and a 135M LM for generative regularization with unlabeled raw-text initialization, ElicitR consistently improves BRIGHT by 16-29% relative across 0.1B–3B retriever scales while maintaining performance on BEIR. At 3B, ElicitR reaches an nDCG@10 of 23.1, substantially outperforming larger models trained with far more curated pairs and proprietary APIs. Further analyses show that ElicitR prevents overfitting, improves retrieval calibration, and remains robust to batch sizes, supporting its practicality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single TransformerZhengbao Jiang, Luyu Gao, Zhiruo Wang, Jun Araki 等EMNLP 2022 · 被引用 14 次
- RaDeR: Reasoning-aware Dense Retrieval ModelsDebrup Das, Seán Ó Nualláin, Razieh RahimiEMNLP 2025 · 被引用 1 次
- Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language ModelsOrion Weller, Benjamin Van Durme, Dawn J. Lawrie, Ashwin Paranjape 等ICLR 2025
- Improving Text Embeddings with Large Language ModelsLiang Wang, Nan Yang, Xiaolong Huang, Linjun Yang 等ACL 2024
相关 Paper
- Training Dense Retrievers with Multiple Positive PassagesBenben Wang, Minghao Tang, Hengran Zhang, Jiafeng Guo 等KDD 2026 · 被引用 1 次
- Reason-to-Rank: Distilling Direct and Comparative Reasoning from Large Language Models for Document RerankingYuelyu Ji, Zhuochun Li, Rui Meng, Daqing HeSIGIR 2025 · 被引用 3 次
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen 等ICLR 2026 · 被引用 3 次
- Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls ShortXuan Lu, Haohang Huang, Rui Meng, Yaohui Jin 等ICLR 2026 · 被引用 11 次
- GLEN: Generative Retrieval via Lexical Index LearningSunkyung Lee, Minjin Choi, Jongwuk LeeEMNLP 2023 · 被引用 6 次
