Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Álvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
Abstract
The lack of time-efficient and reliable evaluation methods hamper the development of conversational dialogue systems (chatbots). Evaluations requiring humans to converse with chatbots are time and cost-intensive, put high cognitive demands on the human judges, and yield low-quality results. In this work, we introduce Spot The Bot, a cost-efficient and robust evaluation framework that replaces human-bot conversations with conversations between bots. Human judges then only annotate for each entity in a conversation whether they think it is human or not (assuming there are humans participants in these conversations). These annotations then allow us to rank chatbots regarding their ability to mimic the conversational behavior of humans. Since we expect that all bots are eventually recognized as such, we incorporate a metric that measures which chatbot can uphold human-like behavior the longest, i.e., Survival Analysis. This metric has the ability to correlate a bot's performance to certain of its characteristics (e.g., fluency or sensibleness), yielding interpretable results. The comparably low cost of our framework allows for frequent evaluations of chatbots during their evaluation cycle. We empirically validate our claims by applying Spot The Bot to three domains, evaluating several stateof-the-art chatbots, and drawing comparisons to related work. The framework is released as a ready-to-use tool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22a43a09-9f16-48b7-b9e1-a3f047112ec5Cited by top-tier papers7
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
- ChatMatch: Evaluating Chatbots by Autonomous Chat TournamentsRuolan Yang, Zitong Li, Haifeng Tang, Kenny Q. ZhuACL 2022 · 12 citations
- Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue SystemsSarah E. Finch, James D. Finch, Jinho D. ChoiACL 2023 · 10 citations
- Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive SetupSijia Liu, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis et al.AAAI 2023 · 2 citations
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
Builds on2
Related papers
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 22 citations
- Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsTianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu et al.ACL 2022
- Designing Effective Interview Chatbots: Automatic Chatbot Profiling and Design Suggestion Generation for Chatbot DebuggingXu Han, Michelle Zhou, Matthew J. Turner, Tom YehCHI 2021 · 53 citations
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou et al.ACL 2020 · 144 citations
- Active Evaluation: Efficient NLG Evaluation with Few Pairwise ComparisonsAkash Kumar Mohankumar, Mitesh M. KhapraACL 2022 · 8 citations
