ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
Tianjian Liu, Fanqi Wan, Jiajian Guo, Xiaojun Quan
摘要
Proactive dialogue has emerged as a critical and challenging research problem in advancing large language models (LLMs).Existing works predominantly focus on domain-specific or task-oriented scenarios, which leads to fragmented evaluations and limits the comprehensive exploration of models' proactive dialogue abilities.In this work, we propose Proac-tiveEval, a unified framework for evaluating proactive dialogue capabilities of LLMs.This framework decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains.Moreover, it also enables the automatic generation of diverse and challenging evaluation data.Based on the proposed framework, we develop 328 evaluation environments spanning 6 distinct domains.Through experiments with 22 different types of LLMs, we show that DeepSeek-R1 and Claude-3.7-Sonnetexhibit exceptional performance on target planning and dialogue guidance tasks, respectively.Finally, we investigate how reasoning capabilities influence proactive behaviors and discuss their implications for future model development.Our code and data are available at the repository.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMsXiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen 等NeurIPS 2025 · 被引用 63 次
- DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationZeming Liu, Haifeng Wang, Zhengyu Niu, Hua Wu 等EMNLP 2021 · 被引用 39 次
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
- ComPeer: A Generative Conversational Agent for Proactive Peer SupportTianjian Liu, Hongzheng Zhao, Yuheng Liu, Xingbo Wang 等UIST 2024 · 被引用 27 次
相关 Paper
- T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by StepZehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu 等ACL 2024 · 被引用 7 次
- MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li 等ACL 2026 · 被引用 12 次
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in InstructionsXingwei He, Qianru Zhang, Pengfei Chen, Guanhua Chen 等AAAI 2026 · 被引用 2 次
