PMIScore: An Unsupervised Approach to Quantify Dialogue Engagement
Yongkang Guo, Zhihuan Huang, Yuqing Kong
摘要
High dialogue engagement is a crucial indicator of an effective conversation. A reliable measure of engagement could help benchmark large language models, enhance the effectiveness of human-computer interactions, or improve personal communication skills. However, quantifying engagement is challenging, since it is subjective and lacks a ''gold standard''. This paper proposes PMIScore, an efficient unsupervised approach to quantify dialogue engagement. It uses pointwise mutual information (PMI), which is the probability of generating a response conditioning on the conversation history. Thus, PMIScore offers a clear interpretation of engagement. As directly computing PMI is intractable due to the complexity of dialogues, PMIScore learned it through a dual form of divergence. The algorithm includes generating positive and negative dialogue pairs, extracting embeddings by large language models (LLMs), and training a small neural network using a mutual information loss function. We validated PMIScore on both synthetic and real-world datasets. Our results demonstrate the effectiveness of PMIScore in PMI estimation and the reasonableness of the PMI metric itself.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven ConversationHao Zhou, Chujie Zheng, Kaili Huang, Minlie Huang 等ACL 2020 · 被引用 106 次
- Benchmarking LLMs' Judgments with No Gold StandardShengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing KongICLR 2025
相关 Paper
- Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded DialogsYatin Nandwani, Vineet Kumar, Dinesh Raghu, Sachindra Joshi 等EMNLP 2023 · 被引用 4 次
- Linguistic Dependencies and Statistical DependenceJacob Louis Hoover, Wenyu Du, Alessandro Sordoni, Timothy J. O'DonnellEMNLP 2021 · 被引用 2 次
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 被引用 62 次
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao 等AAAI 2026
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
