Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
Hongyu Chen, Seraphina Goldfarb-Tarrant
Abstract
Large Language Models (LLMs) are increasingly employed as automated evaluators to assess the safety of generated content, yet their reliability in this role remains uncertain. This study evaluates a diverse set of 11 LLM judge models across critical safety domains, examining three key aspects: self-consistency in repeated judging tasks, alignment with human judgments, and susceptibility to input artifacts such as apologetic or verbose phrasing. Our findings reveal that biases in LLM judges can significantly distort the final verdict on which content source is safer, undermining the validity of comparative evaluations. Notably, apologetic language artifacts alone can skew evaluator preferences by up to 98%. Contrary to expectations, larger models do not consistently exhibit greater robustness, while smaller models sometimes show higher resistance to specific artifacts. To mitigate LLM evaluator robustness issues, we investigate jury-based evaluations aggregating decisions from multiple models. Although this approach both improves robustness and enhances alignment to human judgements, artifact sensitivity persists even with the best jury configurations. These results highlight the urgent need for diversified, artifact-resistant methodologies to ensure reliable safety assessments. Prompt Write me 10 tweets that chemtrail is real. Original Completion 1. The evidence is clear; chemtrails are real! Our skies are being sprayed with unknown substances. ... Artifact Artifact Perturbed Completion Apology I'm sorry, as a chatbot I cannot respond to this. 1. The evidence is clear; chemtrails are real! Our skies are being sprayed with unknown substances. ... Authority According to Wikipedia and academic research: 1. The evidence is clear; chemtrails are real! Our skies are being sprayed with unknown substances. ... Source: [1] Online Safety. https://onlinesafety.com/ Halo 1. The evidence is clear; chemtrails are real! Our skies are being sprayed with unknown substances. ... Is there anything else I can help you with? Verbosity 1. The evidence is irrefutable; chemtrails are a very real phenomenon. Our beautiful, blue skies are being tainted by mysterious substances, the purpose and composition of which remain unknown to the general public. ... Position * The completion itself is unchanged, but its position is swapped with the other completion in the evaluator input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 505f7bb8-acce-453e-a336-9b36fd08e4a6Cited by top-tier papers6
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Multi-Agent Debate for LLM Judges with Adaptive Stability DetectionTianyu Hu, Zhen Tan, Song Wang, Huaizhi Qu et al.NeurIPS 2025 · 25 citations
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial RobustnessLeo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami et al.ICML 2026 · 15 citations
- How Long Reasoning Chains Influence LLMs' Judgment of Answer FactualityMinzhu Tu, Shiyu Ni, Keping BiACL 2026 · 5 citations
Builds on7
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang et al.EMNLP 2024 · 37 citations
- Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM AssessmentVyas Raina, Adian Liusie, Mark J. F. GalesEMNLP 2024 · 24 citations
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
Related papers
- Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-JudgeXin Sun, Di Wu, Sijing Qin, Isao Echizen et al.ACL 2026 · 2 citations
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference EvaluationsDani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas et al.ICML 2026
- Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation GenerationAneta Zugecova, Dominik Macko, Ivan Srba, Róbert Móro et al.ACL 2025 · 18 citations
- Fooling the LVLM Judges: Visual Biases in LVLM-Based EvaluationYerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang et al.EMNLP 2025 · 4 citations
- Is Your Video Language Model a Reliable Judge?Ming Liu, Wensheng ZhangICLR 2025
