CASN: Class-Aware Score Network for Textual Adversarial Detection
Rong Bao, Rui Zheng, Liang Ding, Qi Zhang, Dacheng Tao
Abstract
Adversarial detection aims to detect adversarial samples that threaten the security of deep neural networks, which is an essential step toward building robust AI systems. Density-based estimation is widely considered as an effective technique by explicitly modeling the distribution of normal data and identifying adversarial ones as outliers. However, these methods suffer from significant performance degradation when the adversarial samples lie close to the non-adversarial data manifold. To address this limitation, we propose a score-based generative method to implicitly model the data distribution. Our approach utilizes the gradient of the log-density data distribution and calculates the distribution gap between adversarial and normal samples through multi-step iterations using Langevin dynamics. In addition, we use supervised contrastive learning to guide the gradient estimation using label information, which avoids collapsing to a single data manifold and better preserves the anisotropy of the different labeled data distributions. Experimental results on three text classification tasks upon four advanced attack algorithms show that our approach is a significant improvement (+15.2 F1 score on average against previous SOTA) over previous detection methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9028c98-a269-45ac-affc-00cdd983c3a6Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
Related papers
- DMHM: Density-aware Manifold Learning and Hybrid Mahalanobis Energy for LLMs-generated Text DetectionTianle Liu, Zhiliang Tian, Zhen Huang, Tianlun Liu et al.ACL 2026
- OSTAR: Optimized Statistical Text-classifier with Adversarial ResistanceYuhan Yao, Feifei Kou, Lei Shi, Xiao Yang et al.NeurIPS 2025
- Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation ScoreShuhai Zhang, Feng Liu, Jiahao Yang, Yifan Yang et al.ICML 2023 · 39 citations
- GAT: Generative Adversarial Training for Adversarial Example Detection and Robust ClassificationXuwang Yin, Soheil Kolouri, Gustavo K. RohdeICLR 2020 · 47 citations
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang et al.ICCV 2021 · 70 citations
