ACL2026

The Mark Fades: Adaptive Evolutionary Paraphrase-based Attack against LLM Watermarks

Yusheng Zhao, Jian Zhao, Tianle Zhang, Feng Wei, Xuelong Li

摘要

While LLM watermarking is pivotal for identifying machine-generated content, existing paraphrase-based attacks struggle to achieve an optimal balance between watermark removal efficacy and the text quality. To address this limitation, We propose TSAPA, a training-free evolutionary framework that formulates watermark removal as a constrained multi-objective optimization problem. By leveraging genetic algorithms to navigate the Pareto front, TSAPA utilizes a Pseudo-Log-Likelihood (PLL)-guided mutation strategy to precisely target and modify watermark-carrying tokens. Extensive experiments on Qwen3 series (1.7B/8B/32B) across diverse watermarking schemes demonstrate that TSAPA achieves an attack success rate (ASR) exceeding 90% while maintaining superior text semantic fidelity, significantly outperforming baseline methods. This work exposes critical vulnerabilities in current watermark techniques and provides a novel perspective for their robust evaluation.