Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
摘要
The lack of time-efficient and reliable evaluation methods hamper the development of conversational dialogue systems (chatbots). Evaluations requiring humans to converse with chatbots are time and cost-intensive, put high cognitive demands on the human judges, and yield low-quality results. In this work, we introduce Spot The Bot, a cost-efficient and robust evaluation framework that replaces human-bot conversations with conversations between bots. Human judges then only annotate for each entity in a conversation whether they think it is human or not (assuming there are humans participants in these conversations). These annotations then allow us to rank chatbots regarding their ability to mimic the conversational behavior of humans. Since we expect that all bots are eventually recognized as such, we incorporate a metric that measures which chatbot can uphold human-like behavior the longest, i.e., Survival Analysis. This metric has the ability to correlate a bot's performance to certain of its characteristics (e.g., fluency or sensibleness), yielding interpretable results. The comparably low cost of our framework allows for frequent evaluations of chatbots during their evaluation cycle. We empirically validate our claims by applying Spot The Bot to three domains, evaluating several stateof-the-art chatbots, and drawing comparisons to related work. The framework is released as a ready-to-use tool.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi 等AAAI 2023 · 被引用 112 次
- ChatMatch: Evaluating Chatbots by Autonomous Chat TournamentsRuolan Yang, Zitong Li, Haifeng Tang, Kenny Q. ZhuACL 2022 · 被引用 12 次
- Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue SystemsSarah E. Finch, James D. Finch, Jinho D. ChoiACL 2023 · 被引用 10 次
- Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive SetupSijia Liu, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis 等AAAI 2023 · 被引用 2 次
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
它引用的顶会 Paper2
相关 Paper
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 被引用 22 次
- Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsTianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu 等ACL 2022
- Designing Effective Interview Chatbots: Automatic Chatbot Profiling and Design Suggestion Generation for Chatbot DebuggingXu Han, Michelle Zhou, Matthew J. Turner, Tom YehCHI 2021 · 被引用 53 次
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou 等ACL 2020 · 被引用 144 次
- Active Evaluation: Efficient NLG Evaluation with Few Pairwise ComparisonsAkash Kumar Mohankumar, Mitesh M. KhapraACL 2022 · 被引用 8 次
