DA³: A Distribution-Aware Adversarial Attack against Language Models
Yibo Wang, Xiangjue Dong, James Caverlee, Philip S. Yu
摘要
Language models can be manipulated by adversarial attacks, which introduce subtle perturbations to input data. While recent attack methods can achieve a relatively high attack success rate (ASR), we've observed that the generated adversarial examples have a different data distribution compared with the original examples. Specifically, these adversarial examples exhibit reduced confidence levels and greater divergence from the training data distribution. Consequently, they are easy to detect using straightforward detection methods, diminishing the efficacy of such attacks. To address this issue, we propose a Distribution-Aware Adversarial Attack (DA 3 ) method. DA 3 considers the distribution shifts of adversarial examples to improve attacks' effectiveness under detection methods. We further design a novel evaluation metric, the Non-detectable Attack Success Rate (NASR), which integrates both ASR and detectability for the attack task. We conduct experiments on four widely used datasets to validate the attack effectiveness and transferability of adversarial examples generated by DA 3 against both the white-box BERT-BASE and ROBERTA-BASE models and the black-box LLAMA2-7B model 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 被引用 789 次
相关 Paper
- RAFT: Realistic Attacks to Fool Text DetectorsJames Wang, Ran Li, Junfeng Yang, Chengzhi MaoEMNLP 2024 · 被引用 2 次
- DSRM: Boost Textual Adversarial Training with Distribution Shift Risk MinimizationSongyang Gao, Shihan Dou, Yan Liu, Xiao Wang 等ACL 2023 · 被引用 3 次
- Language Model Detectors Are Easily Optimized AgainstCharlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma 等ICLR 2024 · 被引用 18 次
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
- Generating Distributional Adversarial Examples to Evade Statistical DetectorsYigitcan Kaya, Muhammad Bilal Zafar, Sergül Aydöre, Nathalie Rauschmayr 等ICML 2022 · 被引用 7 次
