Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation
Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, Kewei Tu
摘要
Open-domain dialogue generation has gained increasing attention in Natural Language Processing. Its evaluation requires a holistic means. Human ratings are deemed as the gold standard. As human evaluation is inefficient and costly, an automated substitute is highly desirable. In this paper, we propose holistic evaluation metrics that capture different aspects of open-domain dialogues. Our metrics consist of (1) GPT-2 based context coherence between sentences in a dialogue, (2) GPT-2 based fluency in phrasing, (3) n-gram based diversity in responses to augmented queries, and (4) textual-entailment-inference based logical self-consistency. The empirical validity of our metrics is demonstrated by strong correlations with human judgments. We open source the code and relevant materials. 1 * Equal contributions
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
- Adversarial Attack and Defense of Structured Prediction ModelsWenjuan Han, Liwen Zhang, Yong Jiang, Kewei TuEMNLP 2020 · 被引用 32 次
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs 等EMNLP 2022 · 被引用 13 次
- ChatMatch: Evaluating Chatbots by Autonomous Chat TournamentsRuolan Yang, Zitong Li, Haifeng Tang, Kenny Q. ZhuACL 2022 · 被引用 12 次
- IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue EvaluationZhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang 等EMNLP 2022 · 被引用 7 次
它引用的顶会 Paper1
相关 Paper
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 被引用 49 次
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin 等EMNLP 2020 · 被引用 73 次
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
- Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesZekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng 等ACL 2021
