Studying the Effects of Cognitive Biases in Evaluation of Conversational Agents
Sashank Santhanam, Alireza Karduni, Samira Shaikh
摘要
Humans quite frequently interact with conversational agents. The rapid advancement in generative language modeling through neural networks has helped advance the creation of intelligent conversational agents. Researchers typically evaluate the output of their models through crowdsourced judgments, but there are no established best practices for conducting such studies. Moreover, it is unclear if cognitive biases in decision-making are affecting crowdsourced workers' judgments when they undertake these tasks. To investigate, we conducted a between-subjects study with 77 crowdsourced workers to understand the role of cognitive biases, specifically anchoring bias, when humans are asked to evaluate the output of conversational agents. Our results provide insight into how best to evaluate conversational agents. We find increased consistency in ratings across two experimental conditions may be a result of anchoring bias. We also determine that external factors such as time and prior experience in similar tasks have effects on inter-rater consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- "If I Had All the Time in the World": Ophthalmologists' Perceptions of Anchoring Bias Mitigation in Clinical AI SupportAnne Kathrine Petersen Bach, Trine Munch Nørgaard, Jens Christian Brok, Niels van BerkelCHI 2023 · 被引用 46 次
- How Accurate Does It Feel? - Human Perception of Different Types of Classification MistakesAndrea Papenmeier, Dagmar Kern, Daniel Hienert, Yvonne Kammerer 等CHI 2022 · 被引用 25 次
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati 等ACL 2024 · 被引用 19 次
- Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMsClemencia Siro, Mohammad Aliannejadi, Maarten de RijkeSIGIR 2024 · 被引用 3 次
- Metacognitive Demands and Strategies While Using Off-The-Shelf AI Conversational Agents for Health Information SeekingShri Harini Ramesh, Foroozan Daneshzand, Babak Rashidi, Shriti Raj 等CHI 2026 · 被引用 1 次
相关 Paper
- Deciding Fast and Slow: The Role of Cognitive Biases in AI-assisted Decision-makingCharvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R. Varshney 等CSCW 2022 · 被引用 184 次
- LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical StudyNan Zhuang, Boyu Cao, Yi Yang, Jing Xu 等AAAI 2025 · 被引用 4 次
- Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational AgentsStephen Pilli, Vivek NallurCHI 2026 · 被引用 2 次
- Improving Worker Engagement Through Conversational Microtask CrowdsourcingSihang Qiu, Ujwal Gadiraju, Alessandro BozzonCHI 2020 · 被引用 66 次
- AI-Moderated Decision-Making: Capturing and Balancing Anchoring Bias in Sequential Decision TasksJessica Maria Echterhoff, Matin Yarmand, Julian J. McAuleyCHI 2022 · 被引用 29 次
