KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification
Lingfeng Shen, Shoushan Li, Ying Chen
摘要
Recent work has shown that current text classification models are vulnerable to a small adversarial perturbation on inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current adversarial training methods have two principal problems: a drop in model's generalization and ineffective defending against other text attacks. In this paper, we propose a Keywordbias-aware Adversarial Text Generation model (KATG) that implicitly generates adversarial sentences using a generatordiscriminator structure. Instead of using a benign sentence to generate an adversarial sentence, the KATG model utilizes extra multiple benign sentences (namely prior sentences) to guide adversarial sentence generation. Furthermore, to cover more perturbations used in existing attacks, a keyword-biasbased sampling is proposed to select sentences containing biased words as prior sentences. Besides, to effectively utilize prior sentences, a generative flow mechanism is proposed to construct a latent semantic space for learning a latent representation of the prior sentences. Experiments demonstrate that adversarial sentences generated by our KATG model can strengthen the generalization and the robustness of text classification models. Benign Sentence Sixthreezero is good, I've used it for a long time, only changed because I got tired of the same old bike. (Pos) Prior Sentences S1: Blackberry may work on the systems, but I'm not willing to take that chance on a new expensive phone. (Neg) S2: Iphone4s is in ok previously used condition as stated. But I was disappointed I couldn't activate the phone upon arrival. (Neg) Adv. Sentence Amazing Iphone4s, used it for so long , only changed because I got tired of the old expensive Blackberry. (Pos) Table 1: Benign sentence, prior sentences and adversarial sentence used in our KATG model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Trickle-down Impact of Reward Inconsistency on RLHFLingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin 等ICLR 2024 · 被引用 10 次
- TextShield: Beyond Successfully Detecting Adversarial Sentences in text classificationLingfeng Shen, Ze Zhang, Haiyun Jiang, Ying ChenICLR 2023
它引用的顶会 Paper5
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationLibo Qin, Wanxiang Che, Yangming Li, Minheng Ni 等AAAI 2020 · 被引用 100 次
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 被引用 92 次
- Latent Variable Modelling with Hyperbolic Normalizing FlowsAvishek Joey Bose, Ariella Smofsky, Renjie Liao, Prakash Panangaden 等ICML 2020 · 被引用 76 次
- Adversarial Attack and Defense of Structured Prediction ModelsWenjuan Han, Liwen Zhang, Yong Jiang, Kewei TuEMNLP 2020 · 被引用 32 次
相关 Paper
- Precisely the Point: Adversarial Augmentations for Faithful and Informative Text GenerationWenhao Wu, Wei Li, Jiachen Liu, Xinyan Xiao 等EMNLP 2022 · 被引用 4 次
- Don't Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting TextAshim Gupta, Carter Wood Blum, Temma Choji, Yingjie Fei 等ACL 2023 · 被引用 7 次
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution Based Text AttacksXiaosen Wang, Yichen Yang, Yihe Deng, Kun HeAAAI 2021 · 被引用 98 次
- Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningQin Liu, Rui Zheng, Bao Rong, Jingyi Liu 等ACL 2022 · 被引用 35 次
- TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven OptimizationBairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang 等ICLR 2023 · 被引用 1 次
