CASN: Class-Aware Score Network for Textual Adversarial Detection
Rong Bao, Rui Zheng, Liang Ding, Qi Zhang, Dacheng Tao
摘要
Adversarial detection aims to detect adversarial samples that threaten the security of deep neural networks, which is an essential step toward building robust AI systems. Density-based estimation is widely considered as an effective technique by explicitly modeling the distribution of normal data and identifying adversarial ones as outliers. However, these methods suffer from significant performance degradation when the adversarial samples lie close to the non-adversarial data manifold. To address this limitation, we propose a score-based generative method to implicitly model the data distribution. Our approach utilizes the gradient of the log-density data distribution and calculates the distribution gap between adversarial and normal samples through multi-step iterations using Langevin dynamics. In addition, we use supervised contrastive learning to guide the gradient estimation using label information, which avoids collapsing to a single data manifold and better preserves the anisotropy of the different labeled data distributions. Experimental results on three text classification tasks upon four advanced attack algorithms show that our approach is a significant improvement (+15.2 F1 score on average against previous SOTA) over previous detection methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
相关 Paper
- DMHM: Density-aware Manifold Learning and Hybrid Mahalanobis Energy for LLMs-generated Text DetectionTianle Liu, Zhiliang Tian, Zhen Huang, Tianlun Liu 等ACL 2026
- OSTAR: Optimized Statistical Text-classifier with Adversarial ResistanceYuhan Yao, Feifei Kou, Lei Shi, Xiao Yang 等NeurIPS 2025
- Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation ScoreShuhai Zhang, Feng Liu, Jiahao Yang, Yifan Yang 等ICML 2023 · 被引用 39 次
- GAT: Generative Adversarial Training for Adversarial Example Detection and Robust ClassificationXuwang Yin, Soheil Kolouri, Gustavo K. RohdeICLR 2020 · 被引用 47 次
- Adversarial Example Detection Using Latent Neighborhood GraphAhmed Abusnaina, Yuhang Wu, Sunpreet S. Arora, Yizhen Wang 等ICCV 2021 · 被引用 70 次
