Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems
Sarah E. Finch, James D. Finch, Jinho D. Choi
摘要
Despite tremendous advancements in dialogue systems, stable evaluation still requires human judgments producing notoriously high-variance metrics due to their inherent subjectivity.Moreover, methods and labels in dialogue evaluation are not fully standardized, especially for open-domain chats, with a lack of work to compare and assess the validity of those approaches.The use of inconsistent evaluation can misinform the performance of a dialogue system, which becomes a major hurdle to enhance it.Thus, a dimensional evaluation of chat-oriented open-domain dialogue systems that reliably measures several aspects of dialogue capabilities is desired.This paper presents a novel human evaluation method to estimate the rates of manypasted macro 'LN' dialogue system behaviors.Our method is used to evaluate four state-of-the-art open-domain dialogue systems and compared with existing approaches.The analysis demonstrates that our behavior method is more suitable than alternative Likert-style or comparative approaches for dimensional evaluation of these systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational SearchHideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P. de Vries 等SIGIR 2024 · 被引用 24 次
- Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen 等ACL 2026 · 被引用 7 次
- Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention RewardsMing Li, Pei Chen, Zhenhao Zhang, Tao Yang 等ACL 2026 · 被引用 3 次
- SCBench: A KV Cache-Centric Analysis of Long-Context MethodsYucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo 等ICLR 2025
它引用的顶会 Paper24
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 被引用 329 次
- CEM: Commonsense-Aware Empathetic Response GenerationSahand Sabour, Chujie Zheng, Minlie HuangAAAI 2022 · 被引用 196 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck 等ACL 2020 · 被引用 120 次
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
相关 Paper
- Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsTianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu 等ACL 2022
- MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationChen Zhang, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou LiAAAI 2022 · 被引用 22 次
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
- FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationChen Zhang, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs 等EMNLP 2022 · 被引用 13 次
- Just Adjust One Prompt: Enhancing In-Context Dialogue Scoring via Constructing the Optimal Subgraph of Demonstrations and PromptsJiashu Pu, Ling Cheng, Lu Fan, Tangjie Lv 等EMNLP 2023 · 被引用 2 次
