Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue Evaluation
Peiwen Yuan, Xinglin Wang, Jiayi Shi, Bin Sun, Yiwei Li
摘要
Turn-level dialogue evaluation models (TDEMs), using self-supervised learning (SSL) framework, have achieved state-of-the-art performance in open-domain dialogue evaluation. However, these models inevitably face two potential problems. First, they have low correlations with humans on medium coherence samples as the SSL framework often brings training data with unbalanced coherence distribution. Second, the SSL framework leads TDEM to nonuniform score distribution. There is a danger that the nonuniform score distribution will weaken the robustness of TDEM through our theoretical analysis. To tackle these problems, we propose B etter C orrelation and R obustness (BCR), a distribution-balanced self-supervised learning framework for TDEM. Given a dialogue dataset, BCR offers an effective training set reconstructing method to provide coherence-balanced training signals and further facilitate balanced evaluating abilities of TDEM. To get a uniform score distribution, a novel loss function is proposed, which can adjust adaptively according to the uniformity of score distribution estimated by kernel density estimation. Comprehensive experiments on 17 benchmark datasets show that vanilla BERT-base using BCR outperforms SOTA methods significantly by 11.3% on average. BCR also demonstrates strong generalization ability as it can lead multiple SOTA methods to attain better correlation and robustness. Code and datasets: https://github.com/ypw0102/Better-Correlation-and-Robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- BatchEval: Towards Human-like Text EvaluationPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang 等ACL 2024
- Focused Large Language Models are Stable Many-Shot LearnersPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang 等EMNLP 2024
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin 等EMNLP 2020 · 被引用 73 次
相关 Paper
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs 等EMNLP 2022 · 被引用 13 次
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog EvaluationWeixin Liang, James Zou, Zhou YuACL 2020 · 被引用 25 次
- Rating Distribution Calibration for Selection Bias Mitigation in RecommendationsHaochen Liu, Da Tang, Ji Yang, Xiangyu Zhao 等WWW 2022 · 被引用 31 次
- DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank UtterancesXiaodong Gu, Kang Min Yoo, Jung-Woo HaAAAI 2021 · 被引用 83 次
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang 等ACL 2021
