KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification
Lingfeng Shen, Shoushan Li, Ying Chen
Abstract
Recent work has shown that current text classification models are vulnerable to a small adversarial perturbation on inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current adversarial training methods have two principal problems: a drop in model's generalization and ineffective defending against other text attacks. In this paper, we propose a Keywordbias-aware Adversarial Text Generation model (KATG) that implicitly generates adversarial sentences using a generatordiscriminator structure. Instead of using a benign sentence to generate an adversarial sentence, the KATG model utilizes extra multiple benign sentences (namely prior sentences) to guide adversarial sentence generation. Furthermore, to cover more perturbations used in existing attacks, a keyword-biasbased sampling is proposed to select sentences containing biased words as prior sentences. Besides, to effectively utilize prior sentences, a generative flow mechanism is proposed to construct a latent semantic space for learning a latent representation of the prior sentences. Experiments demonstrate that adversarial sentences generated by our KATG model can strengthen the generalization and the robustness of text classification models. Benign Sentence Sixthreezero is good, I've used it for a long time, only changed because I got tired of the same old bike. (Pos) Prior Sentences S1: Blackberry may work on the systems, but I'm not willing to take that chance on a new expensive phone. (Neg) S2: Iphone4s is in ok previously used condition as stated. But I was disappointed I couldn't activate the phone upon arrival. (Neg) Adv. Sentence Amazing Iphone4s, used it for so long , only changed because I got tired of the old expensive Blackberry. (Pos) Table 1: Benign sentence, prior sentences and adversarial sentence used in our KATG model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext daf0d0dc-0c24-4f09-81d8-1752b37ef863Cited by top-tier papers2
- The Trickle-down Impact of Reward Inconsistency on RLHFLingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin et al.ICLR 2024 · 10 citations
- TextShield: Beyond Successfully Detecting Adversarial Sentences in text classificationLingfeng Shen, Ze Zhang, Haiyun Jiang, Ying ChenICLR 2023
Builds on5
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationLibo Qin, Wanxiang Che, Yangming Li, Minheng Ni et al.AAAI 2020 · 100 citations
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 92 citations
- Latent Variable Modelling with Hyperbolic Normalizing FlowsAvishek Joey Bose, Ariella Smofsky, Renjie Liao, Prakash Panangaden et al.ICML 2020 · 76 citations
- Adversarial Attack and Defense of Structured Prediction ModelsWenjuan Han, Liwen Zhang, Yong Jiang, Kewei TuEMNLP 2020 · 32 citations
Related papers
- Precisely the Point: Adversarial Augmentations for Faithful and Informative Text GenerationWenhao Wu, Wei Li, Jiachen Liu, Xinyan Xiao et al.EMNLP 2022 · 4 citations
- Don't Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting TextAshim Gupta, Carter Wood Blum, Temma Choji, Yingjie Fei et al.ACL 2023 · 7 citations
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution Based Text AttacksXiaosen Wang, Yichen Yang, Yihe Deng, Kun HeAAAI 2021 · 98 citations
- Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningQin Liu, Rui Zheng, Bao Rong, Jingyi Liu et al.ACL 2022 · 35 citations
- TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven OptimizationBairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang et al.ICLR 2023 · 1 citation
