ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
Tianjian Liu, Fanqi Wan, Jiajian Guo, Xiaojun Quan
Abstract
Proactive dialogue has emerged as a critical and challenging research problem in advancing large language models (LLMs).Existing works predominantly focus on domain-specific or task-oriented scenarios, which leads to fragmented evaluations and limits the comprehensive exploration of models' proactive dialogue abilities.In this work, we propose Proac-tiveEval, a unified framework for evaluating proactive dialogue capabilities of LLMs.This framework decomposes proactive dialogue into target planning and dialogue guidance, establishing evaluation metrics across various domains.Moreover, it also enables the automatic generation of diverse and challenging evaluation data.Based on the proposed framework, we develop 328 evaluation environments spanning 6 distinct domains.Through experiments with 22 different types of LLMs, we show that DeepSeek-R1 and Claude-3.7-Sonnetexhibit exceptional performance on target planning and dialogue guidance tasks, respectively.Finally, we investigate how reasoning capabilities influence proactive behaviors and discuss their implications for future model development.Our code and data are available at the repository.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98808db0-b1c1-455a-95fb-97d3b932b3a3Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 · 86 citations
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMsXiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen et al.NeurIPS 2025 · 63 citations
- DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationZeming Liu, Haifeng Wang, Zhengyu Niu, Hua Wu et al.EMNLP 2021 · 39 citations
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
- ComPeer: A Generative Conversational Agent for Proactive Peer SupportTianjian Liu, Hongzheng Zhao, Yuheng Liu, Xingbo Wang et al.UIST 2024 · 27 citations
Related papers
- T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by StepZehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu et al.ACL 2024 · 7 citations
- MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li et al.ACL 2026 · 12 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in InstructionsXingwei He, Qianru Zhang, Pengfei Chen, Guanhua Chen et al.AAAI 2026 · 2 citations
