Grounded in Reality: Learning and Deploying Proactive LLM from Offline Logs
Fei Wei, Daoyuan Chen, Ce Wang, Yilun Huang, Yushuo Chen, Xuchen Pan, Yaliang Li, Bolin Ding
摘要
Large Language Models (LLMs) excel as passive responders, but teaching them to be proactive, goal-oriented partners-a critical capability in high-stakes domains-remains a major challenge. Current paradigms either myopically optimize single-turn attributes or rely on brittle, high-cost user simulators, creating a persistent "reality gap". To bridge this gap, we introduce Learn-to-Ask, a general, simulator-free framework for learning and deploying proactive dialogue agents directly from offline expert data, bypassing the need to model complex user dynamics. Our key insight is to reframe the offline policy learning problem by leveraging the observed future of each expert trajectory. This allows us to infer a dense, turn-by-turn reward signal grounded in the expert's revealed strategy, decomposing the intractable long-horizon problem into a series of supervised learning tasks, and training a policy to output a structured (action, state assessment) tuple, governing both what to ask and, crucially, when to stop. To ensure reward fidelity, our Automated Grader Calibration pipeline systematically purges noise from the LLM-based reward model with minimal human supervision. Empirically, we demonstrate the efficacy of Learn-to-Ask in a real-world medical dataset, using LLMs of varying sizes up to 32B. Our approach culminates in the successful deployment of LLMs into a live, large-scale online AI service. In rigorous in-house evaluations, our model was launched and achieved performance even superior to human experts, proving our framework's ability to translate offline data into tangible, real-world impact. We hope this work provides a practical and economically viable blueprint for transforming passive LLMs into proactive, goal-oriented LLM applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine 等ICML 2024 · 被引用 163 次
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 被引用 112 次
- Generating Multi-turn Clarification for Web Information SeekingZiliang Zhao, Zhicheng DouWWW 2024 · 被引用 15 次
相关 Paper
- Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement LearningYunghwei Lai, Kaiming Liu, Ziyue Wang, Weizhi Ma 等ICLR 2026 · 被引用 10 次
- Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded ReasoningJiayuan Zhu, Jiazhen Pan, Yuyuan Liu, Fenglin Liu 等EMNLP 2025 · 被引用 1 次
- Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question AnsweringXueren Ge, Sahil Murtaza, Anthony Cortez, Homa AlemzadehAAAI 2026 · 被引用 2 次
- Ask and Retrieve Knowledge: Towards Proactive Asking with Imperfect Information in Medical Multi-turn DialoguesBolin Zhang, Shengwei Wang, Yangqin Jiang, Dianbo Sui 等SIGIR 2025
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
