Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual Information
Kun Zhao, Bohao Yang, Chenghua Lin, Wenge Rong, Aline Villavicencio, Xiaohui Cui
Abstract
The long-standing one-to-many issue of the open-domain dialogues poses significant challenges for automatic evaluation methods, i.e., there may be multiple suitable responses which differ in semantics for a given conversational context.To tackle this challenge, we propose a novel learning-based automatic evaluation metric (CMN), which can robustly evaluate open-domain dialogues by augmenting Conditional Variational Autoencoders (CVAEs) with a Next Sentence Prediction (NSP) objective and employing Mutual Information (MI) to model the semantic similarity of text in the latent space. Experimental results on two open-domain dialogue datasets demonstrate the superiority of our method compared with a wide range of baselines, especially in handling responses which are distant to the "golden" reference responses in semantics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecda2a4b-5f60-4861-924a-c9c8cbfabe48Cited by top-tier papers3
- Uncertainty Regularized Evidential RegressionKai Ye, Tiejin Chen, Hua Wei, Liang ZhanAAAI 2024 · 15 citations
- Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation PapersChen Tang, Shun Wang, Tomas Goldsack, Chenghua LinEMNLP 2023 · 5 citations
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningWeifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li et al.EMNLP 2023 · 2 citations
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent SpaceChunyuan Li, Xiang Gao, Yuan Li, Baolin Peng et al.EMNLP 2020 · 132 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog EvaluationWeixin Liang, James Zou, Zhou YuACL 2020 · 25 citations
- Generating Dialogue Responses from a Semantic Latent SpaceWei-Jen Ko, Avik Ray, Yilin Shen, Hongxia JinEMNLP 2020 · 4 citations
- RADE: Reference-Assisted Dialogue Evaluation for Open-Domain DialogueZhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang et al.ACL 2023 · 5 citations
- Speaker Sensitive Response Evaluation ModelJinYeong Bak, Alice OhACL 2020 · 10 citations
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 49 citations
