Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification
Chenhao Dang, Jing Ma
Abstract
A persistent challenge in text classification (TC) is that enhancing model robustness against adversarial attacks typically degrades performance on clean data. We argue that this challenge can be resolved by modeling the distribution of clean samples in the encoder's embedding manifold. To this end, we propose the Manifold-Correcting Causal Flow (M C 2 F ), a two-module system that operates directly on sentence embeddings. A Stratified Riemannian Continuous Normalizing Flow (SR-CNF) learns the density of the clean data manifold. It identifies out-of-distribution embeddings, which are then corrected by a Geodesic Purification Solver. This solver projects adversarial points back onto the learned manifold via the shortest path, restoring a clean, semantically coherent representation. We conducted extensive evaluations on text classification (TC) across three datasets and multiple adversarial attacks. The results demonstrate that our method, M C 2 F , not only establishes a new state-of-the-art in adversarial robustness but also fully preserves performance on clean data, even yielding modest gains in Accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- Fisher SAM: Information Geometry and Sharpness Aware MinimisationMinyoung Kim, Da Li, Shell Xu Hu, Timothy M. HospedalesICML 2022 · 96 citations
Related papers
- Textual Manifold-based Defense Against Natural Language Adversarial ExamplesDang Minh Nguyen, Anh Tuan LuuEMNLP 2022 · 13 citations
- Semantic Perturbations with Normalizing Flows for Improved GeneralizationOguz Kaan Yüksel, Sebastian U. Stich, Martin Jaggi, Tatjana ChavdarovaICCV 2021 · 13 citations
- Adversarial Purification with the Manifold HypothesisZhaoyuan Yang, Zhiwei Xu, Jing Zhang, Richard I. Hartley et al.AAAI 2024 · 12 citations
- Tractable Density Estimation on Learned Manifolds with Conformal Embedding FlowsBrendan Leigh Ross, Jesse C. CresswellNeurIPS 2021 · 39 citations
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu et al.EMNLP 2025
