Lune

ICLR2026顶会

Q-RAG: Long Context Multi‑Step Retrieval via Value‑Based Embedder Training

Artyom Y. Sorokin, Nazar Buzun, Alexander Anokhin, Egor Vedernikov, Petr Anokhin, Mikhail Burtsev, Evgeny Burnaev

2026年份
4被引次数

摘要

Retrieval-Augmented Generation (RAG) methods enhance LLM performance by efficiently filtering relevant context for LLMs, reducing hallucinations and inference cost. However, most existing RAG methods focus on single-step retrieval, which is often insufficient for answering complex questions that require multi-step search. Recently, multi-step retrieval approaches have emerged, typically involving the fine-tuning of small LLMs to perform multi-step retrieval. This type of fine-tuning is highly resource-intensive and does not enable the use of larger LLMs. In this work, we propose Q-RAG, a novel approach that fine-tunes the Embedder model for multi-step retrieval using reinforcement learning (RL). Q-RAG offers a competitive, resource-efficient alternative to existing multi-step retrieval methods for open-domain question answering and achieves state-of-the-art results on the popular long-context benchmarks BabiLong and RULER for contexts up to 10M tokens. Code is available at: https://github.com/griver/Q-RAG.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext e2305ac8-c4e5-47ba-ace1-a1448cb255cc

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖