Lune

ICML2026Top-tier venue

A Game-Theoretic Analysis of Attacks on Large Language Models via Compositional Skills

Xinbo Wu, Huan Zhang, Abhishek Umrawal, Lav Varshney

2026Year

Abstract

Warning: This paper contains potentially offensive or harmful text examples. As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we introduce a theoretical framework that formalizes a game between an attacker hiding its intent via compositional skills and a defender. Within this framework, we design a theoretical best-response attack strategy and show that it is closely related to many existing adversarial prompting methods. We further analyze the resulting game, characterize its equilibria, and reveal inherent advantages for the attacker. Drawing on our theoretical analysis, we also derive a provably optimal defense strategy. Empirically, we evaluate a practical instantiation of the theoretically optimal attack and observe stronger performance relative to existing adversarial prompting approaches in diverse settings encompassing different LLMs and benchmarks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext df41925d-1baf-4fc0-9b73-bc4ff4520053

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines