ChatMatch: Evaluating Chatbots by Autonomous Chat Tournaments
Ruolan Yang, Zitong Li, Haifeng Tang, Kenny Q. Zhu
Abstract
Existing automatic evaluation systems of chatbots mostly rely on static chat scripts as ground truth, which is hard to obtain, and requires access to the models of the bots as a form of “white-box testing”. Interactive evaluation mitigates this problem but requires human involvement. In our work, we propose an interactive chatbot evaluation framework in which chatbots compete with each other like in a sports tournament, using flexible scoring metrics. This framework can efficiently rank chatbots independently from their model architectures and the domains for which they are trained.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0832c39-df9b-4ac2-8b7c-2c182d8cb7d8Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck et al.ACL 2020 · 120 citations
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston et al.ACL 2020 · 18 citations
- Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue SystemsJan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos et al.EMNLP 2020 · 12 citations
Related papers
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu et al.ACL 2025 · 34 citations
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 22 citations
- Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching EffectivenessYuhang Zhou, Xutian Chen, Yixin Cao, Yuchen Ni et al.ICLR 2026
- Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation ApproachHaoming Jiang, Bo Dai, Mengjiao Yang, Tuo Zhao et al.EMNLP 2021 · 5 citations
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildBill Yuchen Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander et al.ICLR 2025
