DA³: A Distribution-Aware Adversarial Attack against Language Models
Yibo Wang, Xiangjue Dong, James Caverlee, Philip S. Yu
Abstract
Language models can be manipulated by adversarial attacks, which introduce subtle perturbations to input data. While recent attack methods can achieve a relatively high attack success rate (ASR), we've observed that the generated adversarial examples have a different data distribution compared with the original examples. Specifically, these adversarial examples exhibit reduced confidence levels and greater divergence from the training data distribution. Consequently, they are easy to detect using straightforward detection methods, diminishing the efficacy of such attacks. To address this issue, we propose a Distribution-Aware Adversarial Attack (DA 3 ) method. DA 3 considers the distribution shifts of adversarial examples to improve attacks' effectiveness under detection methods. We further design a novel evaluation metric, the Non-detectable Attack Success Rate (NASR), which integrates both ASR and detectability for the attack task. We conduct experiments on four widely used datasets to validate the attack effectiveness and transferability of adversarial examples generated by DA 3 against both the white-box BERT-BASE and ROBERTA-BASE models and the black-box LLAMA2-7B model 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95fb8f2f-5a51-44ba-87df-456a26f8543cBuilds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
Related papers
- RAFT: Realistic Attacks to Fool Text DetectorsJames Wang, Ran Li, Junfeng Yang, Chengzhi MaoEMNLP 2024 · 2 citations
- DSRM: Boost Textual Adversarial Training with Distribution Shift Risk MinimizationSongyang Gao, Shihan Dou, Yan Liu, Xiao Wang et al.ACL 2023 · 3 citations
- Language Model Detectors Are Easily Optimized AgainstCharlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma et al.ICLR 2024 · 18 citations
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
- Generating Distributional Adversarial Examples to Evade Statistical DetectorsYigitcan Kaya, Muhammad Bilal Zafar, Sergül Aydöre, Nathalie Rauschmayr et al.ICML 2022 · 7 citations
