Studying the Effects of Cognitive Biases in Evaluation of Conversational Agents
Sashank Santhanam, Alireza Karduni, Samira Shaikh
Abstract
Humans quite frequently interact with conversational agents. The rapid advancement in generative language modeling through neural networks has helped advance the creation of intelligent conversational agents. Researchers typically evaluate the output of their models through crowdsourced judgments, but there are no established best practices for conducting such studies. Moreover, it is unclear if cognitive biases in decision-making are affecting crowdsourced workers' judgments when they undertake these tasks. To investigate, we conducted a between-subjects study with 77 crowdsourced workers to understand the role of cognitive biases, specifically anchoring bias, when humans are asked to evaluate the output of conversational agents. Our results provide insight into how best to evaluate conversational agents. We find increased consistency in ratings across two experimental conditions may be a result of anchoring bias. We also determine that external factors such as time and prior experience in similar tasks have effects on inter-rater consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e78fbfc-4e1b-4e33-8a97-8047bc67147aCited by top-tier papers6
- "If I Had All the Time in the World": Ophthalmologists' Perceptions of Anchoring Bias Mitigation in Clinical AI SupportAnne Kathrine Petersen Bach, Trine Munch Nørgaard, Jens Christian Brok, Niels van BerkelCHI 2023 · 46 citations
- How Accurate Does It Feel? - Human Perception of Different Types of Classification MistakesAndrea Papenmeier, Dagmar Kern, Daniel Hienert, Yvonne Kammerer et al.CHI 2022 · 25 citations
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
- Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMsClemencia Siro, Mohammad Aliannejadi, Maarten de RijkeSIGIR 2024 · 3 citations
- Metacognitive Demands and Strategies While Using Off-The-Shelf AI Conversational Agents for Health Information SeekingShri Harini Ramesh, Foroozan Daneshzand, Babak Rashidi, Shriti Raj et al.CHI 2026 · 1 citation
Related papers
- Deciding Fast and Slow: The Role of Cognitive Biases in AI-assisted Decision-makingCharvi Rastogi, Yunfeng Zhang, Dennis Wei, Kush R. Varshney et al.CSCW 2022 · 184 citations
- LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical StudyNan Zhuang, Boyu Cao, Yi Yang, Jing Xu et al.AAAI 2025 · 4 citations
- Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational AgentsStephen Pilli, Vivek NallurCHI 2026 · 2 citations
- Improving Worker Engagement Through Conversational Microtask CrowdsourcingSihang Qiu, Ujwal Gadiraju, Alessandro BozzonCHI 2020 · 66 citations
- AI-Moderated Decision-Making: Capturing and Balancing Anchoring Bias in Sequential Decision TasksJessica Maria Echterhoff, Matin Yarmand, Julian J. McAuleyCHI 2022 · 29 citations
