Revisiting Character-level Adversarial Attacks for Language Models
Elías Abad-Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
摘要
Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While character-level attacks easily maintain semantics, they have received less attention as they cannot easily adopt popular gradient-based methods, and are thought to be easy to defend. Challenging these beliefs, we introduce Charmer, an efficient query-based adversarial attack capable of achieving high attack success rate (ASR) while generating highly similar adversarial examples. Our method successfully targets both small (BERT) and large (Llama 2) models. Specifically, on BERT with SST-2, Charmer improves the ASR in 4.84% points and the USE similarity in 8% points with respect to the previous art. Our implementation is available in https://github.com/LIONS-EPFL/Charmer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu 等NeurIPS 2025 · 被引用 4 次
- SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak AttacksSeungwon Jeong, Jiwoo Jeong, Hyeonjin Kim, Yunseok Lee 等ICLR 2026 · 被引用 2 次
- Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word ImportanceBryan E. Tuck, Rakesh M. VermaAAAI 2026 · 被引用 1 次
- Certified Robustness Under Bounded Levenshtein DistanceElías Abad-Rocamora, Grigorios Chrysos, Volkan CevherICLR 2025
- FlipAttack: Jailbreak LLMs via FlippingYue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu 等ICML 2025
它引用的顶会 Paper14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski 等USENIX Security 2019 · 被引用 466 次
相关 Paper
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu 等ACL 2020 · 被引用 188 次
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui 等ICLR 2024 · 被引用 146 次
- DA³: A Distribution-Aware Adversarial Attack against Language ModelsYibo Wang, Xiangjue Dong, James Caverlee, Philip S. YuEMNLP 2024 · 被引用 2 次
- Character-level White-Box Adversarial Attacks against Transformers via Attachable Subwords SubstitutionAiwei Liu, Honghai Yu, Xuming Hu, Shu'ang Li 等EMNLP 2022 · 被引用 20 次
- A Strong Baseline for Query Efficient Attacks in a Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiEMNLP 2021 · 被引用 38 次
