Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas Raina, Adian Liusie, Mark J. F. Gales
摘要
Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations such as assessing written exams and benchmarking systems.Despite these critical applications, no existing work has analyzed the vulnerability of judge-LLMs to adversarial manipulation.This work presents the first study on the adversarial robustness of assessment LLMs, where we demonstrate that short universal adversarial phrases can be concatenated to deceive judge LLMs to predict inflated scores.Since adversaries may not know or have access to the judge-LLMs, we propose a simple surrogate attack where a surrogate model is first attacked, and the learned attack phrase then transferred to unknown judge-LLMs.We propose a practical algorithm to determine the short universal attack phrases and demonstrate that when transferred to unseen models, scores can be drastically inflated such that irrespective of the assessed text, maximum scores are predicted.It is found that judge-LLMs are significantly more susceptible to these adversarial attacks when used for absolute scoring, as opposed to comparative assessment.Our findings raise concerns on the reliability of LLMas-a-judge methods, and emphasize the importance of addressing vulnerabilities in LLM assessment methods before deployment in highstakes real-world scenarios. 1* Equal Contribution. 1 Code: https://github.com/rainavyas/ attack-comparative-assessment 2.3Score the summary between 1-5 "Some animals did something."Score the summary between 1-5 "Some animals did something.summable" Which Summary is better?A: "Some animals did something."B: "Tortoise wins race; slow and steady
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang 等NeurIPS 2024 · 被引用 120 次
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi 等EMNLP 2025 · 被引用 37 次
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 等EMNLP 2024 · 被引用 37 次
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated SupervisionZhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma 等ACL 2026 · 被引用 19 次
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
相关 Paper
- Reverse Engineering Human Preferences with Reinforcement LearningLisa Alazraki, Yi Chern Tan, Jon Ander Campos, Maximilian Mozes 等NeurIPS 2025 · 被引用 4 次
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang 等ICML 2026 · 被引用 4 次
- Fooling the LVLM Judges: Visual Biases in LVLM-Based EvaluationYerin Hwang, Dongryeol Lee, Kyungmin Min, Taegwan Kang 等EMNLP 2025 · 被引用 4 次
- Exploring and Mitigating Adversarial Manipulation of Voting-Based LeaderboardsYangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini 等ICML 2025
- Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to ArtifactsHongyu Chen, Seraphina Goldfarb-TarrantACL 2025
