ChatMatch: Evaluating Chatbots by Autonomous Chat Tournaments
Ruolan Yang, Zitong Li, Haifeng Tang, Kenny Q. Zhu
2022年份
12被引次数
摘要
Existing automatic evaluation systems of chatbots mostly rely on static chat scripts as ground truth, which is hard to obtain, and requires access to the models of the bots as a form of “white-box testing”. Interactive evaluation mitigates this problem but requires human involvement. In our work, we propose an interactive chatbot evaluation framework in which chatbots compete with each other like in a sports tournament, using flexible scoring metrics. This framework can efficiently rank chatbots independently from their model architectures and the domains for which they are trained.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck 等ACL 2020 · 被引用 120 次
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou 等ACL 2020 · 被引用 69 次
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston 等ACL 2020 · 被引用 18 次
- Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos 等EMNLP 2020 · 被引用 12 次
相关 Paper
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 等ACL 2025 · 被引用 34 次
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 被引用 22 次
- Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching EffectivenessYuhang Zhou, Xutian Chen, Yixin Cao, Yuchen Ni 等ICLR 2026
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation ApproachHaoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao 等EMNLP 2021 · 被引用 5 次
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildBill Yuchen Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander 等ICLR 2025
