RaDeR: Reasoning-aware Dense Retrieval Models
Debrup Das, Seán Ó Nualláin, Razieh Rahimi
摘要
We propose RaDeR, a set of reasoning-based dense retrieval models trained with data derived from mathematical problem solving using large language models (LLMs). Our method leverages retrieval-augmented reasoning trajectories of an LLM and self-reflective relevance evaluation, enabling the creation of both diverse and hard-negative samples for reasoning-intensive relevance. RaDeR retrievers, trained for mathematical reasoning, effectively generalize to diverse reasoning tasks in the BRIGHT and RAR-b benchmarks, consistently outperforming strong baselines in overall performance. Notably, RaDeR achieves significantly higher performance than baselines on the Math and Coding splits. In addition, RaDeR presents the first dense retriever that outperforms BM25 when queries are Chain-of-Thought reasoning steps, underscoring the critical role of reasoning-based retrieval to augment reasoning language models. Furthermore, RaDeR achieves comparable or superior performance while using only 2.5% of the training data used by the concurrent work REASONIR, highlighting the quality of our synthesized training data. Our code, data, and retrieval models are publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document RetrievalJianlyu Chen, Junwei Lan, Chaofan Li, Defu Lian 等ACL 2026 · 被引用 13 次
- Inferential Question AnsweringJamshid Mozafari, Hamed Zamani, Guido Zuccon, Adam JatowtWWW 2026
- Beyond Markovian Forgetfulness: Episodic Memory for Reasoning-Intensive RetrievalDohyeon Lee, Yeonseok Jeong, Seung-won HwangACL 2026
- Internalizing Explicit Reasoning into Latent Space for Dense RetrievalJiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long 等SIGIR 2026
- ElicitR: Unlocking Latent Reasoning in Dense Retrievers via Generative RegularizationFengyu Cai, Iryna Gurevych, Heinz KoepplICML 2026
它引用的顶会 Paper14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
- Find Tailored Step Example for Next Step: a Targeted Step-wise Retrieval Framework for Guiding LLM ReasoningCheng Yang, Zhenya Huang, Liyang He, Weibo Gao 等KDD 2026
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM GuidanceSijia Yao, Pengcheng Huang, Zhenghao Liu, Yu Gu 等EMNLP 2025 · 被引用 6 次
- MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle TaskYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang 等ICLR 2026 · 被引用 5 次
- Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive TasksMinki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi 等NeurIPS 2023 · 被引用 128 次
