GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, Xiaodan Liang
Abstract
Automatically evaluating dialogue coherence is a challenging but high-demand ability for developing high-quality open-domain dialogue systems. However, current evaluation metrics consider only surface features or utterancelevel semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows. Here, we first consider that the graph structure constituted with topics in a dialogue can accurately depict the underlying communication logic, which is a more natural way to produce persuasive metrics. Capitalized on the topic-level dialogue graph, we propose a new evaluation metric GRADE, which stands for Graph-enhanced Representations for Automatic Dialogue Evaluation. Specifically, GRADE incorporates both coarsegrained utterance-level contextualized representations and fine-grained topic-level graph representations to evaluate dialogue coherence. The graph representations are obtained by reasoning over topic-level dialogue graphs enhanced with the evidence from a commonsense graph, including k-hop neighboring representations and hop-attention weights. Experimental results show that our GRADE significantly outperforms other state-of-the-art metrics on measuring diverse dialogue models in terms of the Pearson and Spearman correlations with human judgements. Besides, we release a new large-scale human evaluation benchmark to facilitate future research on automatic metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e907c8e2-9f0e-4ed3-ab53-a68863ed24c5Cited by top-tier papers27
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao et al.EMNLP 2022 · 103 citations
- Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response GenerationPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim et al.ACL 2022 · 45 citations
- InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri et al.EMNLP 2022 · 26 citations
- Ditch the Gold Standard: Re-evaluating Conversational Question AnsweringHuihan Li, Tianyu Gao, Manan Goenka, Danqi ChenACL 2022 · 23 citations
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 22 citations
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Dialogues Are Not Just Text: Modeling Cognition for Dialogue Coherence EvaluationXue Li, Jia Su, Yang Yang, Zipeng Gao et al.AAAI 2024 · 5 citations
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang et al.ACL 2021
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act FlowsJianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji et al.EMNLP 2022 · 4 citations
- Improving Knowledge-Aware Dialogue Generation via Knowledge Base Question AnsweringJian Wang, Junhao Liu, Wei Bi, Xiaojiang Liu et al.AAAI 2020 · 52 citations
