Revisiting Character-level Adversarial Attacks for Language Models
Elías Abad-Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
Abstract
Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While character-level attacks easily maintain semantics, they have received less attention as they cannot easily adopt popular gradient-based methods, and are thought to be easy to defend. Challenging these beliefs, we introduce Charmer, an efficient query-based adversarial attack capable of achieving high attack success rate (ASR) while generating highly similar adversarial examples. Our method successfully targets both small (BERT) and large (Llama 2) models. Specifically, on BERT with SST-2, Charmer improves the ASR in 4.84% points and the USE similarity in 8% points with respect to the previous art. Our implementation is available in https://github.com/LIONS-EPFL/Charmer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b153101-0c92-49c2-8e2f-70e20bd69f83Cited by top-tier papers10
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu et al.NeurIPS 2025 · 4 citations
- SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak AttacksSeungwon Jeong, Jiwoo Jeong, Hyeonjin Kim, Yunseok Lee et al.ICLR 2026 · 2 citations
- Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word ImportanceBryan E. Tuck, Rakesh M. VermaAAAI 2026 · 1 citation
- Certified Robustness Under Bounded Levenshtein DistanceElías Abad-Rocamora, Grigorios Chrysos, Volkan CevherICLR 2025
- FlipAttack: Jailbreak LLMs via FlippingYue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu et al.ICML 2025
Builds on14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski et al.USENIX Security 2019 · 466 citations
Related papers
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
- DA³: A Distribution-Aware Adversarial Attack against Language ModelsYibo Wang, Xiangjue Dong, James Caverlee, Philip S. YuEMNLP 2024 · 2 citations
- Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords SubstitutionAiwei Liu, Honghai Yu, Xuming Hu, Shu'ang Li et al.EMNLP 2022 · 20 citations
- A Strong Baseline for Query Efficient Attacks in a Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiEMNLP 2021 · 38 citations
