Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog Evaluation
Weixin Liang, James Zou, Zhou Yu
摘要
Open Domain dialog system evaluation is one of the most important challenges in dialog research. Existing automatic evaluation metrics, such as BLEU are mostly referencebased. They calculate the difference between the generated response and a limited number of available references. Likert-score based self-reported user rating is widely adopted by social conversational systems, such as Amazon Alexa Prize chatbots. However, selfreported user rating suffers from bias and variance among different users. To alleviate this problem, we formulate dialog evaluation as a comparison task. We also propose an automatic evaluation model CMADE (Comparison Model for Automatic Dialog Evaluation) that automatically cleans self-reported user ratings as it trains on them. Specifically, we first use a self-supervised method to learn better dialog feature representation, and then use KNN and Shapley to remove confusing samples. Our experiments show that CMADE achieves 89.2% accuracy in the dialog comparison task. Our implementation is available at https://github.com/Weixin-Liang/ dialog_evaluation_CMADE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training ConflictsWeixin Liang, James ZouICLR 2022 · 被引用 103 次
- Do Response Selection Models Really Know What's Next? Utterance Manipulation Strategies for Multi-turn Response SelectionTaesun Whang, Dongyub Lee, Dongsuk Oh, Chanhee Lee 等AAAI 2021 · 被引用 70 次
- STEAM: Self-Supervised Taxonomy Expansion with Mini-PathsYue Yu, Yinghao Li, Jiaming Shen, Hao Feng 等KDD 2020 · 被引用 47 次
- ALICE: Active Learning with Contrastive Natural Language ExplanationsWeixin Liang, James Zou, Zhou YuEMNLP 2020 · 被引用 36 次
- DU-Shapley: A Shapley Value Proxy for Efficient Dataset ValuationFelipe Garrido-Lucero, Benjamin Heymann, Maxime Vono, Patrick Loiseau 等NeurIPS 2024 · 被引用 19 次
它引用的顶会 Paper2
相关 Paper
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 被引用 49 次
- Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual InformationKun Zhao, Bohao Yang, Chenghua Lin, Wenge Rong 等ACL 2023 · 被引用 14 次
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 被引用 22 次
- RADE: Reference-Assisted Dialogue Evaluation for Open-Domain DialogueZhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang 等ACL 2023 · 被引用 5 次
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
