ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
Jianlyu Chen, Junwei Lan, Chaofan Li, Defu Lian, Zheng Liu
摘要
In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes three key technical contributions. First, we propose ReMixer, a new data synthesis method that overcomes the triviality problem prevalent in previous synthetic datasets, enabling large-scale production of 82K highquality training samples. Second, we design Redapter, a self-adaptive learning algorithm that dynamically adjusts training each sample's weight based on its reasoning intensity. This allows the model to effectively capture the complex semantic relationships between queries and documents. Third, we implement Rea-sonEmbed across multiple backbones of varying sizes, all of which achieve superior performance on reasoning-intensive retrieval tasks. Notably, our ReasonEmbed-Qwen3-8B model offers a record-high nDCG@10 score of 38.1 on the BRIGHT benchmark (SU et al., 2025) , which significantly outperforms existing text embedding models. We will fully open-source our created resources in ReasonEmbed to push forward the research advancement in this field 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Internalizing Explicit Reasoning into Latent Space for Dense RetrievalJiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long 等SIGIR 2026
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
- With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind SpotsZeinab Taghavi, Ali Modarressi, Hinrich Schuetze, Andreas MarfurtICML 2026
它引用的顶会 Paper11
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- ReasonRank: Empowering Passage Ranking with Strong Reasoning AbilityWenhan Liu, Xinyu Ma, Weiwei Sun, Yutao Zhu 等ACL 2026 · 被引用 43 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
- Retro*: Optimizing LLMs for Reasoning-Intensive Document RetrievalJunwei Lan, Jianlyu Chen, Zheng Liu, Chaofan Li 等ICLR 2026 · 被引用 8 次
- RaDeR: Reasoning-aware Dense Retrieval ModelsDebrup Das, Seán Ó Nualláin, Razieh RahimiEMNLP 2025 · 被引用 1 次
相关 Paper
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi 等ICLR 2025
- Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search SystemsYilun Zhao, Jinbiao Wei, Tingyu Song, Siyue Zhang 等ACL 2026
- ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained AlignmentHao Yang, Yifan Ji, Zhipeng Xu, Zhenghao Liu 等SIGIR 2026 · 被引用 4 次
- ElicitR: Unlocking Latent Reasoning in Dense Retrievers via Generative RegularizationFengyu Cai, Iryna Gurevych, Heinz KoepplICML 2026
- RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image RetrievalYijiang Li, Kunal Kotian, Ali Marjaninejad, Meir Friedenberg 等CVPR 2026
