DEAM: Dialogue Coherence Evaluation using AMR-based Semantic Manipulations
Sarik Ghazarian, Nuan Wen, Aram Galstyan, Nanyun Peng
Abstract
Automatic evaluation metrics are essential for the rapid development of open-domain dialogue systems as they facilitate hyperparameter tuning and comparison between models. Although recently proposed trainable conversation-level metrics have shown encouraging results, the quality of the metrics is strongly dependent on the quality of training data. Prior works mainly resort to heuristic text-level manipulations (e.g. utterances shuffling) to bootstrap incoherent conversations (negative examples) from coherent dialogues (positive examples). Such approaches are insufficient to appropriately reflect the incoherence that occurs in interactions between advanced dialogue models and humans. To tackle this problem, we propose DEAM, a Dialogue coherence Evaluation metric that relies on Abstract Meaning Representation (AMR) to apply semantic-level Manipulations for incoherent (negative) data generation. AMRs naturally facilitate the injection of various types of incoherence sources, such as coreference inconsistency, irrelevancy, contradictions, and decrease engagement, at the semantic level, thus resulting in more natural incoherent samples. Our experiments show that DEAM achieves higher correlations with human judgments compared to baseline methods on several dialog datasets by significant margins. We also show that DEAM can distinguish between coherent and incoherent dialogues generated by baseline manipulations, whereas those baseline models cannot detect incoherent examples generated by DEAM. Our results demonstrate the potential of AMR-based semantic manipulations for natural negative example generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81d0446a-0646-4c4a-aa74-cb32302c18b5Cited by top-tier papers11
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang et al.AAAI 2024 · 57 citations
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal et al.ACL 2024 · 30 citations
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao et al.ACL 2023 · 20 citations
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs et al.EMNLP 2022 · 13 citations
- A Survey of AMR ApplicationsShira Wein, Juri OpitzEMNLP 2024 · 7 citations
Builds on8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Diverse and Informative Dialogue Generation with Context-Specific Commonsense Knowledge AwarenessSixing Wu, Ying Li, Dawei Zhang, Yang Zhou et al.ACL 2020 · 104 citations
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin et al.EMNLP 2020 · 73 citations
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 62 citations
- Improving AMR Parsing with Sequence-to-Sequence Pre-trainingDongqin Xu, Junhui Li, Muhua Zhu, Min Zhang et al.EMNLP 2020 · 57 citations
Related papers
- Semantic Representation for Dialogue ModelingXuefeng Bai, Yulong Chen, Linfeng Song, Yue ZhangACL 2021
- Dialogues Are Not Just Text: Modeling Cognition for Dialogue Coherence EvaluationXue Li, Jia Su, Yang Yang, Zipeng Gao et al.AAAI 2024 · 5 citations
- Towards Quantifiable Dialogue Coherence EvaluationZheng Ye, Liucun Lu, Lishan Huang, Liang Lin et al.ACL 2021
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
