Lune

EMNLP2022顶会

IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation

Zhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang, Xin Miao

2022年份
7被引次数

摘要

Evaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE. However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality or several qualities. To mitigate the problem, this paper proposes an interpretable, multi-faceted, and controllable framework IM 2 (Interpretable and M ulti-category Integrated M etric) to combine a large number of metrics which are good at measuring different qualities. The IM 2 framework first divides current popular dialogue qualities into different categories and then applies or proposes dialogue metrics to measure the qualities within each category and finally generates an overall IM 2 score. An initial version of IM 2 was submitted to the AAAI 2022 Track5.1@DSTC10 challenge 1 and took the 2 nd place on both of the development and test leaderboard. After the competition, we develop more metrics and improve the performance of our model. We compare IM 2 with other 13 current dialogue metrics and experimental results show that IM 2 correlates more strongly with human judgments than any of them on each evaluated dataset 2 .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f8ce05bc-5f3a-4cff-a73b-bfaf212916cd

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖