A Human-machine Collaborative Framework for Evaluating Malevolence in Dialogues
Yangjun Zhang, Pengjie Ren, Maarten de Rijke
摘要
Conversational dialogue systems (CDSs) are hard to evaluate due to the complexity of natural language. Automatic evaluation of dialogues often shows insufficient correlation with human judgements. Human evaluation is reliable but labor-intensive. We introduce a humanmachine collaborative framework, HMCEval, that can guarantee reliability of the evaluation outcomes with reduced human effort. HMCEval casts dialogue evaluation as a sample assignment problem, where we need to decide to assign a sample to a human or a machine for evaluation. HMCEval includes a model confidence estimation module to estimate the confidence of the predicted sample assignment, and a human effort estimation module to estimate the human effort should the sample be assigned to human evaluation, as well as a sample assignment execution module that finds the optimum assignment solution based on the estimated confidence and effort. We assess the performance of HMCEval on the task of evaluating malevolence in dialogues. The experimental results show that HMCEval achieves around 99% evaluation accuracy with half of the human effort spared, showing that HMCEval provides reliable evaluation outcomes while reducing human effort by a large amount.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long ContextsYifei Yu, Qian-Wen Zhang, Lingfeng Qiao, Di Yin 等EMNLP 2025 · 被引用 2 次
- Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration ApproachChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsTianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu 等ACL 2022
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang 等ACL 2021
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation ApproachHaoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao 等EMNLP 2021 · 被引用 5 次
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 被引用 22 次
- Improving Multi-label Malevolence Detection in Dialogues through Multi-faceted Label Correlation EnhancementYangjun Zhang, Pengjie Ren, Wentao Deng, Zhumin Chen 等ACL 2022 · 被引用 10 次
