ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
Simon Lupart, Mohammad Aliannejadi, Evangelos Kanoulas
摘要
We present ChatR1, a reasoning framework based on reinforcement learning (RL) for conversational question answering (CQA). Reasoning plays an important role in CQA, where user intent evolves across dialogue turns, and utterances are often underspecified, requiring contextual interpretation, query reformulation, and dynamic coordination between retrieval and generation. Unlike static 'rewrite, retrieve, and generate' pipelines, ChatR1 interleaves search and reasoning across turns, enabling exploratory and adaptive behaviors learned through RL. To address the challenge of sparse and delayed rewards in RL, we propose an intent-aware reward that provides turn-level feedback by aligning retrieval and reasoning with evolving user goals. ChatR1 demonstrates strong performance on both 3B and 7B model backbones, outperforming competitive models on five CQA datasets, measured by different metrics (F1, BERTScore, and LLMas-judge). We include a diverse set of CQA datasets to cover topic shifts, evolving intents, mixed-initiative dialogues, and multi-document grounding, testing ChatR1's performance from various aspects. Ablation studies confirm the effectiveness of the intent-aware reward. Our analyses further reveal diverse reasoning trajectories and effective use of the search tool. ChatR1 also generalizes robustly across domains, demonstrating that RL-based reasoning enables more flexible and context-aware behavior than static CQA pipelines. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- ChatQA: Surpassing GPT-4 on Conversational QA and RAGZihan Liu, Wei Ping, Rajarshi Roy, Peng Xu 等NeurIPS 2024 · 被引用 121 次
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu 等SIGIR 2020 · 被引用 84 次
- Few-Shot Conversational Dense RetrievalShi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng 等SIGIR 2021 · 被引用 75 次
- MultiDoc2Dial: Modeling Dialogues Grounded in Multiple DocumentsSong Feng, Siva Sankalp Patel, Hui Wan, Sachindra JoshiEMNLP 2021 · 被引用 42 次
相关 Paper
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter 等EMNLP 2022 · 被引用 37 次
- DVCQR: Dual-View Conversational Query Rewriting with Stage-wise Reinforcement LearningChenyi Li, Xinhui Tu, Zaixiang WangACL 2026
- Learning Contextual Retrieval for Robust Conversational SearchSeunghan Yang, Juntae Lee, Jihwan Bang, Kyuhong Shim 等EMNLP 2025 · 被引用 3 次
- CRAF: A Clinical Reasoning-Adaptive Framework via Reinforcement Learning for Similar Case RetrievalJie Lin, Lei Jiang, Zongyi Chen, Liansheng WangAAAI 2026
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo 等EMNLP 2024 · 被引用 7 次
