Sampling-aware Adversarial Attacks Against Large Language Models
Tim Beyer, Yan Scholten, Leo Schwinn, Stephan Günnemann
Abstract
To guarantee safe and robust deployment of large language models (LLMs) at scale, it is critical to accurately assess their adversarial robustness. Existing adversarial attacks typically target harmful responses in single-point greedy generations, overlooking the inherently stochastic nature of LLMs and overestimating robustness. We show that for the goal of eliciting harmful responses, repeated sampling of model outputs during the attack complements prompt optimization and serves as a strong and efficient attack vector. By casting attacks as a resource allocation problem between optimization and sampling, we empirically determine compute-optimal trade-offs and show that integrating sampling into existing attacks boosts success rates by up to 37% and improves efficiency by up to two orders of magnitude. We further analyze how distributions of output harmfulness evolve during an adversarial attack, discovering that many common optimization strategies have little effect on output harmfulness. Finally, we introduce a label-free proof-of-concept objective based on entropy maximization, demonstrating how our sampling-aware perspective enables new optimization targets. Overall, our findings establish the importance of sampling in attacks to accurately assess and strengthen LLM safety at scale.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial RobustnessLeo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami et al.ICML 2026 · 15 citations
- Estimating Tail Risks in Language Model Output DistributionsRico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu et al.ICML 2026 · 3 citations
Related papers
- REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic ObjectiveSimon Geisler, Tom Wollschläger, M. H. I. Abdalla, Vincent Cohen-Addad et al.ICML 2025
- Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe SamplingYiran Zhao, Wenyue Zheng, Tianle Cai, Do Xuan Long et al.NeurIPS 2024 · 46 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
- Semantic Representation Attack against Aligned Large Language ModelsJiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang et al.NeurIPS 2025 · 3 citations
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMsCsaba Dékány, Stefan Balauca, Dimitar I. Dimitrov, Robin Staab et al.NeurIPS 2025 · 12 citations
