Dialogue-RAG: Enhancing Retrieval for LLMs via Node-Linking Utterance Rewriting
Qiwei Li, Teng Xiao, Zuchao Li, Ping Wang, Mengjia Shen, Hai Zhao
Abstract
Large Language Models (LLMs) and Retrieval Augmented Generation (RAG) methods have demonstrated significant potential on tasks across multiple domains. However, ellipses and coreferences, as common phenomena in dialogue scenes, pose challenges to LLMs' understanding and RAG's retrieval accuracy. The previous works ignore the negative impact of this fuzzy data on RAG system. We explore the capabilities of LLMs and RAG systems in dialogue scenarios and use Incomplete Utterance Rewriting (IUR) to complete the key information in dialogue to enhance retrieval. Besides, we propose a lightweight IUR model for query rewriting. It is an end-to-end framework for node linking and iterative inference, incorporating two newly proposed probing semantic features derived from generative pre-training. This framework treats IUR as a series of link decisions on the input sequence and the incrementally constructed rewriting outputs. To test the performance of RAG system in the model multi-round dialogue scenario, we construct an RAG dialogue dataset on English and Chinese, Dialogue-RAG-MULTI-v1.0. Experiment results show that utterance rewriting can effectively improve the retrieval and generation ability of RAG system in dialogue scenes. Experiments on IUR tasks demonstrate the excellent performance of our lightweight IUR method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0ea9821-9d0a-4e67-a6fc-0e9da577d948Cited by top-tier papers2
- End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question AnsweringJiliang Hu, Zuchao Li, Baoyuan Qi, Guoming Liu et al.AAAI 2026
- From Parameters to Performance: A Data-Driven Study on LLM Structure and DevelopmentSuqing Wang, Zuchao Li, Luohe Shi, Bo Du et al.EMNLP 2025
Builds on15
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
Related papers
- Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance AugmentationZhiyu Cao, Peifeng Li, Yaxin Fan, Qiaoming ZhuEMNLP 2024
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement LearningHaoran Luo, Haihong E, Guanting Chen, Qika Lin et al.ICML 2026 · 50 citations
- Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented GenerationShicheng Xu, Liang Pang, Mo Yu, Fandong Meng et al.ACL 2024
- ICR: Iterative Clarification and Rewriting for Conversational SearchZhiyu Cao, Peifeng Li, Qiaoming ZhuEMNLP 2025
- IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner MonologuesDiji Yang, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo et al.SIGIR 2024 · 45 citations
