Humans or LLMs as the Judge? A Study on Judgement Bias
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, Benyou Wang
摘要
Adopting human and large language models (LLM) as judges (a.k.a human-and LLM-as-ajudge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation results. In this paper, we propose a novel framework that is free from referencing groundtruth annotations for investigating Misinformation Oversight Bias, Gender Bias, Authority Bias and Beauty Bias on LLM and human judges. We curate a dataset referring to the revised Bloom's Taxonomy and conduct thousands of evaluations. Results show that human and LLM judges are vulnerable to perturbations to various degrees, and that even the cutting-edge judges possess considerable biases. We further exploit these biases to conduct attacks on LLM judges. We hope that our work can notify the community of the bias and vulnerability of human-and LLMas-a-judge, as well as the urgency of developing robust evaluation systems 1 . Warning: we provide illustrative attack protocols to reveal the vulnerabilities of LLM judges, aiming to develop more robust ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li 等ICLR 2026 · 被引用 74 次
- Search Arena: Analyzing Search-Augmented LLMsMihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li 等ICLR 2026 · 被引用 32 次
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu 等NeurIPS 2025 · 被引用 25 次
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle 等ICLR 2026 · 被引用 24 次
- BLEUBERI: BLEU is a surprisingly effective reward for instruction followingYapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh 等NeurIPS 2025 · 被引用 21 次
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng 等ICLR 2024 · 被引用 299 次
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 被引用 244 次
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
相关 Paper
- Justice or Prejudice? Quantifying Biases in LLM-as-a-JudgeJiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen 等ICLR 2025
- "Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English TextRishav Hada, Agrima Seth, Harshita Diddee, Kalika BaliEMNLP 2023 · 被引用 10 次
- BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge EvaluationPeng Lai, Zhihao Ou, Yong Wang, Longyue Wang 等ICLR 2026 · 被引用 15 次
- Understanding Large Language Model Vulnerabilities to Social Bias AttacksJiaxu Zhao, Meng Fang, Fanghua Ye, Ke Xu 等ACL 2025
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 被引用 270 次
