Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue Systems
Sarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun Peng
Abstract
User engagement is a critical metric for evaluating the quality of open-domain dialogue systems. Prior work has focused on conversation-level engagement by using heuristically constructed features such as the number of turns and the total time of the conversation. In this paper, we investigate the possibility and efficacy of estimating utterance-level engagement and define a novel metric, predictive engagement, for automatic evaluation of open-domain dialogue systems. Our experiments demonstrate that (1) human annotators have high agreement on assessing utterance-level engagement scores; (2) conversation-level engagement scores can be predicted from properly aggregated utterance-level engagement scores. Furthermore, we show that the utterance-level engagement scores can be learned from data. These scores can be incorporated into automatic evaluation metrics for open-domain dialogue systems to improve the correlation with human judgements. This suggests that predictive engagement can be used as a real-time feedback for training better dialogue models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c424101-3f95-48db-8980-9551de1ab9f5Cited by top-tier papers12
- The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare ApplicationsMina Valizadeh, Natalie PardeACL 2022 · 55 citations
- IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue EvaluationZhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang et al.EMNLP 2022 · 7 citations
- RADE: Reference-Assisted Dialogue Evaluation for Open-Domain DialogueZhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang et al.ACL 2023 · 5 citations
- Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue EvaluationPeiwen Yuan, Xinglin Wang, Jiayi Shi, Bin Sun et al.NeurIPS 2023 · 5 citations
- FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act FlowsJianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji et al.EMNLP 2022 · 4 citations
Related papers
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 10 citations
- HERALD: An Annotation Efficient Method to Detect User Disengagement in Social ConversationsWeixin Liang, Kaihui Liang, Zhou YuACL 2021
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett et al.EMNLP 2020 · 67 citations
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs et al.EMNLP 2022 · 13 citations
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
