Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke
摘要
In ad-hoc retrieval, evaluation relies heavily on user actions, including implicit feedback. In a conversational setting such signals are usually unavailable due to the nature of the interactions, and, instead, the evaluation often relies on crowdsourced evaluation labels. The role of user feedback in annotators' assessment of turns in a conversational perception has been little studied. We focus on how the evaluation of task oriented dialogue systems (TDSs), is affected by considering user feedback, explicit or implicit, as provided through the follow-up utterance of a turn being evaluated. We explore and compare two methodologies for assessing TDSs: one includes the user's follow-up utterance and one without. We use both crowdworkers and large language models (LLMs) as annotators to assess system responses across four aspects: relevance, usefulness, interestingness, and explanation quality. Our findings indicate that there is a distinct difference in ratings assigned by both annotator groups in the two setups, indicating that user feedback does influence system evaluation. Workers are more susceptible to user feedback on usefulness and interestingness compared to LLMs on interestingness and relevance. User feedback leads to a more personalized assessment of usefulness by workers, aligning closely with the user's explicit feedback. Additionally, in cases of ambiguous or complex user requests, user feedback improves agreement among crowdworkers. These findings emphasize the significance of user feedback in refining system evaluations and suggest the potential for automated feedback integration in future research. We publicly release the annotated data. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Is GPT-3 a Good Data Annotator?Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia 等ACL 2023 · 被引用 133 次
- Measuring Recommendation Explanation Quality: The Conflicting Goals of ExplanationsKrisztian Balog, Filip RadlinskiSIGIR 2020 · 被引用 89 次
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 被引用 80 次
- Studying the Effects of Cognitive Biases in Evaluation of Conversational AgentsSashank Santhanam, Alireza Karduni, Samira ShaikhCHI 2020 · 被引用 18 次
相关 Paper
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 被引用 153 次
- Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language ModelsHritik Bansal, John Dang, Aditya GroverICLR 2024 · 被引用 28 次
- Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsSarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun PengAAAI 2020 · 被引用 62 次
- Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist DataMalik Khadar, Daniel Runningen, Julia Tang, Stevie Chancellor 等CSCW 2025 · 被引用 2 次
- Interpretable User Satisfaction Estimation for Conversational Systems with Large Language ModelsYing-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang 等ACL 2024 · 被引用 11 次
