LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical Study
Nan Zhuang, Boyu Cao, Yi Yang, Jing Xu, Mingda Xu, Yuxiao Wang, Qi Liu
Abstract
With Large Language Model (LLM) agents taking on more evaluation responsibilities in decision-making, it is essential to recognize their possible biases to guarantee fair and trustworthy AI-supported decisions. This study is the first to thoroughly examine the choice-supportive bias in LLM agents, a cognitive bias that is known to impact human decisionmaking and evaluation. We conduct experiments across 19 open/unopen-source LLM models in five scenarios at maximum, employing both memory-based and evaluation-based tasks adapted and redesigned from human cognitive studies. Our findings show that LLM agents may exhibit biased attribution or evaluation that supports their initial choices, and such bias may persist even if contextual hallucination is not observable. Key findings show that bias manifestation can differ greatly depending on prompt construction and context preservation, and the bias may be mitigated in larger models. Significantly, we observe that the bias increases when the agents perceive they are in control. Our extensive study involving 284 well-educated humans shows that, despite bias, certain LLM agents can still perform better than humans in similar evaluation tasks. This research contributes to the growing area of AI psychology, and the findings underscore the importance of addressing cognitive biases in LLM Agent systems, with wide-ranging implications spanning from improving AI-assisted decision-making to advancing AI safety and ethics. * These authors contributed equally. All equal-contributing authors worked jointly from the first day until the last day, sharing equal workload and contributions. N. Zhuang, B. Cao, M. Xu were responsible for designing and conducting the experiments, Y. Yang was responsible for coding, J. Xu performed the human study, Y. Wang and Q. Liu were responsible for reviewing and supervising.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsLing Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao et al.NeurIPS 2024 · 144 citations
- Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachWeiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan et al.NeurIPS 2024 · 122 citations
Related papers
- Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational AgentsStephen Pilli, Vivek NallurCHI 2026 · 2 citations
- In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM GenerationsMohammad Aflah Khan, Mahsa Amani, Soumi Das, Bishwamittra Ghosh et al.ICLR 2026 · 4 citations
- Exploring Prosocial Irrationality for LLM Agents: A Social Cognition ViewXuan Liu, Jie Zhang, Haoyang Shang, Song Guo et al.ICLR 2025
- Cognitive Biases in LLM-Assisted Software DevelopmentXinyi Zhou, Zeinadsadat Saghi, Sadra Sabouri, Rahul Pandita et al.ICSE 2026
- From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent InteractionsJiayi Li, Xiao Liu, Yansong FengAAAI 2026 · 3 citations
