Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive Setup
Sijia Liu, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis, Di Jin, Andrew Wirth, Yang Liu, Dilek Hakkani-Tur
摘要
Evaluating open-domain conversation models has been an open challenge due to the open-ended nature of conversations. In addition to static evaluations, recent work has started to explore a variety of per-turn and per-dialog interactive evaluation mechanisms and provide advice on the best setup. In this work, we adopt the interactive evaluation framework and further apply to multiple models with a focus on per-turn evaluation techniques. Apart from the widely used setting where participants select the best response among different candidates at each turn, one more novel per-turn evaluation setting is adopted, where participants can select all appropriate responses with different fallback strategies to continue the conversation when no response is selected. We evaluate these settings based on sensitivity and consistency using four GPT2-based models that differ in model sizes or fine-tuning data. To better generalize to any model groups with no prior assumptions on their rankings and control evaluation costs for all setups, we also propose a methodology to estimate the required sample size given a minimum performance gap of interest before running most experiments. Our comprehensive human evaluation results shed light on how to conduct credible human evaluations of open domain dialog systems using the interactive setup, and suggest additional future directions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett 等EMNLP 2020 · 被引用 67 次
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston 等ACL 2020 · 被引用 18 次
- Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos 等EMNLP 2020 · 被引用 12 次
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
- Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsTianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu 等ACL 2022
相关 Paper
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou 等ACL 2020 · 被引用 69 次
- Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual InformationKun Zhao, Bohao Yang, Chenghua Lin, Wenge Rong 等ACL 2023 · 被引用 14 次
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 被引用 49 次
- Speaker Sensitive Response Evaluation ModelJinYeong Bak, Alice OhACL 2020 · 被引用 10 次
- MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang 等EMNLP 2024 · 被引用 14 次
