IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation
Zhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang, Xin Miao
Abstract
Evaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE. However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality or several qualities. To mitigate the problem, this paper proposes an interpretable, multi-faceted, and controllable framework IM 2 (Interpretable and M ulti-category Integrated M etric) to combine a large number of metrics which are good at measuring different qualities. The IM 2 framework first divides current popular dialogue qualities into different categories and then applies or proposes dialogue metrics to measure the qualities within each category and finally generates an overall IM 2 score. An initial version of IM 2 was submitted to the AAAI 2022 Track5.1@DSTC10 challenge 1 and took the 2 nd place on both of the development and test leaderboard. After the competition, we develop more metrics and improve the performance of our model. We compare IM 2 with other 13 current dialogue metrics and experimental results show that IM 2 correlates more strongly with human judgments than any of them on each evaluated dataset 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8ce05bc-5f3a-4cff-a73b-bfaf212916cdBuilds on6
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin et al.EMNLP 2020 · 73 citations
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 62 citations
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 10 citations
- Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesZekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng et al.ACL 2021
Related papers
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs et al.EMNLP 2022 · 13 citations
- Towards Quantifiable Dialogue Coherence EvaluationZheng Ye, Liucun Lu, Lishan Huang, Liang Lin et al.ACL 2021
- Beyond Correlation: Interpretable Evaluation of Machine Translation MetricsStefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba et al.EMNLP 2024 · 1 citation
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 22 citations
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
