ACL2026

Measuring Watermarking under Jailbreaking: ASR Inflation and Goal-Compliance Mismatch

Sungwoo Han, Sangjun Moon, Jingun Kwon, Hidetaka Kamigaito, Manabu Okumura

摘要

Recently, watermarking has attracted growing attention as a practical technique for source attribution of machine-generated text. However, most prior work studies watermarking under benign prompts, while its behavior under jailbreaking prompts remains underexplored. This gap matters because jailbreaking can bypass safety policies and shift the generation regime, raising concerns that watermarking may interact with model alignment under attack. To address this gap, we evaluate six watermarking methods on four LLMs across two jailbreak benchmarks and three settings: Static, Auto-DAN, and DSN. We find that watermarking can inflate judge-based attack success rate, denoted ASR, under jailbreaking, with the largest effects appearing in biased schemes that perturb logits. At the same time, these ASR increases often do not reflect higher harmful-goal compliance when measured by StrongREJECT or by human judgments. This suggests that ASRonly evaluations can be brittle to decoding perturbations and may overestimate harmful-goal compliance, motivating complementary goalcompliance metrics (e.g., StrongREJECT) and human evaluations. WARNING: This paper contains AIgenerated text that is offensive in nature.