Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue Systems
Sarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun Peng
摘要
User engagement is a critical metric for evaluating the quality of open-domain dialogue systems. Prior work has focused on conversation-level engagement by using heuristically constructed features such as the number of turns and the total time of the conversation. In this paper, we investigate the possibility and efficacy of estimating utterance-level engagement and define a novel metric, predictive engagement, for automatic evaluation of open-domain dialogue systems. Our experiments demonstrate that (1) human annotators have high agreement on assessing utterance-level engagement scores; (2) conversation-level engagement scores can be predicted from properly aggregated utterance-level engagement scores. Furthermore, we show that the utterance-level engagement scores can be learned from data. These scores can be incorporated into automatic evaluation metrics for open-domain dialogue systems to improve the correlation with human judgements. This suggests that predictive engagement can be used as a real-time feedback for training better dialogue models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare ApplicationsMina Valizadeh, Natalie PardeACL 2022 · 被引用 55 次
- IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue EvaluationZhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang 等EMNLP 2022 · 被引用 7 次
- RADE: Reference-Assisted Dialogue Evaluation for Open-Domain DialogueZhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang 等ACL 2023 · 被引用 5 次
- Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue EvaluationPeiwen Yuan, Xinglin Wang, Jiayi Shi, Bin Sun 等NeurIPS 2023 · 被引用 5 次
- FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act FlowsJianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji 等EMNLP 2022 · 被引用 4 次
相关 Paper
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
- HERALD: An Annotation Efficient Method to Detect User Disengagement in Social ConversationsWeixin Liang, Kaihui Liang, Zhou YuACL 2021
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett 等EMNLP 2020 · 被引用 67 次
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs 等EMNLP 2022 · 被引用 13 次
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
