Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification
Chenhao Dang, Jing Ma
摘要
A persistent challenge in text classification (TC) is that enhancing model robustness against adversarial attacks typically degrades performance on clean data. We argue that this challenge can be resolved by modeling the distribution of clean samples in the encoder's embedding manifold. To this end, we propose the Manifold-Correcting Causal Flow (M C 2 F ), a two-module system that operates directly on sentence embeddings. A Stratified Riemannian Continuous Normalizing Flow (SR-CNF) learns the density of the clean data manifold. It identifies out-of-distribution embeddings, which are then corrected by a Geodesic Purification Solver. This solver projects adversarial points back onto the learned manifold via the shortest path, restoring a clean, semantically coherent representation. We conducted extensive evaluations on text classification (TC) across three datasets and multiple adversarial attacks. The results demonstrate that our method, M C 2 F , not only establishes a new state-of-the-art in adversarial robustness but also fully preserves performance on clean data, even yielding modest gains in Accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
- Fisher SAM: Information Geometry and Sharpness Aware MinimisationMinyoung Kim, Da Li, Shell Xu Hu, Timothy M. HospedalesICML 2022 · 被引用 96 次
相关 Paper
- Textual Manifold-based Defense Against Natural Language Adversarial ExamplesDang Minh Nguyen, Anh Tuan LuuEMNLP 2022 · 被引用 13 次
- Semantic Perturbations with Normalizing Flows for Improved GeneralizationOguz Kaan Yüksel, Sebastian U. Stich, Martin Jaggi, Tatjana ChavdarovaICCV 2021 · 被引用 13 次
- Adversarial Purification with the Manifold HypothesisZhaoyuan Yang, Zhiwei Xu, Jing Zhang, Richard I. Hartley 等AAAI 2024 · 被引用 12 次
- Tractable Density Estimation on Learned Manifolds with Conformal Embedding FlowsBrendan Leigh Ross, Jesse C. CresswellNeurIPS 2021 · 被引用 39 次
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu 等EMNLP 2025
