PMIScore: An Unsupervised Approach to Quantify Dialogue Engagement
Yongkang Guo, Zhihuan Huang, Yuqing Kong
Abstract
High dialogue engagement is a crucial indicator of an effective conversation. A reliable measure of engagement could help benchmark large language models, enhance the effectiveness of human-computer interactions, or improve personal communication skills. However, quantifying engagement is challenging, since it is subjective and lacks a ''gold standard''. This paper proposes PMIScore, an efficient unsupervised approach to quantify dialogue engagement. It uses pointwise mutual information (PMI), which is the probability of generating a response conditioning on the conversation history. Thus, PMIScore offers a clear interpretation of engagement. As directly computing PMI is intractable due to the complexity of dialogues, PMIScore learned it through a dual form of divergence. The algorithm includes generating positive and negative dialogue pairs, extracting embeddings by large language models (LLMs), and training a small neural network using a mutual information loss function. We validated PMIScore on both synthetic and real-world datasets. Our results demonstrate the effectiveness of PMIScore in PMI estimation and the reasonableness of the PMI metric itself.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95229658-9232-40f5-bf34-99356c532fc0Builds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven ConversationHao Zhou, Chujie Zheng, Kaili Huang, Minlie Huang et al.ACL 2020 · 106 citations
- Benchmarking LLMs' Judgments with No Gold StandardShengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing KongICLR 2025
Related papers
- Pointwise Mutual Information Based Metric and Decoding Strategy for Faithful Generation in Document Grounded DialogsYatin Nandwani, Vineet Kumar, Dinesh Raghu, Sachindra Joshi et al.EMNLP 2023 · 4 citations
- Linguistic Dependencies and Statistical DependenceJacob Louis Hoover, Wenyu Du, Alessandro Sordoni, Timothy J. O'DonnellEMNLP 2021 · 2 citations
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 62 citations
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao et al.AAAI 2026
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
