Certified Robustness Against Natural Language Attacks by Causal Intervention
Haiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng, Hanwang Zhang
摘要
Deep learning models have achieved great success in many fields, yet they are vulnerable to adversarial examples. This paper follows a causal perspective to look into the adversarial vulnerability and proposes Causal Intervention by Semantic Smoothing (CISS), a novel framework towards robustness against natural language attacks. Instead of merely fitting observational data, CISS learns causal effects p(y|do(x)) by smoothing in the latent semantic space to make robust predictions, which scales to deep architectures and avoids tedious construction of noise customized for specific attacks. CISS is provably robust against word substitution attacks, as well as empirically robust even when perturbations are strengthened by unknown attack algorithms. For example, on YELP, CISS surpasses the runner-up by 6.7% in terms of certified robustness against word substitutions, and achieves 79.4% empirical robustness when syntactic attacks are integrated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
- Textual Manifold-based Defense Against Natural Language Adversarial ExamplesDang Minh Nguyen, Anh Tuan LuuEMNLP 2022 · 被引用 13 次
- UniT: A Unified Look at Certified Robust Training against Text Adversarial PerturbationMuchao Ye, Ziyi Yin, Tianrong Zhang, Tianyu Du 等NeurIPS 2023 · 被引用 3 次
- Certified Robustness Under Bounded Levenshtein DistanceElías Abad-Rocamora, Grigorios Chrysos, Volkan CevherICLR 2025
- CertTA: Certified Robustness Made Practical for Learning-Based Traffic AnalysisJinzhu Yan, Zhuotao Liu, Yuyang Xie, Shiyu Liang 等USENIX Security 2025
它引用的顶会 Paper14
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 被引用 533 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 被引用 284 次
- Representation Learning via Invariant Causal MechanismsJovana Mitrovic, Brian McWilliams, Jacob C. Walker, Lars Holger Buesing 等ICLR 2021 · 被引用 281 次
相关 Paper
- Word Level Robustness Enhancement: Fight Perturbation with PerturbationPei Huang, Yuting Yang, Fuqi Jia, Minghao Liu 等AAAI 2022 · 被引用 14 次
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 被引用 8 次
- Adversarial Robustness Through the Lens of CausalityYonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu 等ICLR 2022 · 被引用 65 次
- Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine LearningByung-Kwan Lee, Junho Kim, Yong Man RoICCV 2023 · 被引用 12 次
- Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial ExamplesRuichu Cai, Yuxuan Zhu, Jie Qiao, Zefeng Liang 等AAAI 2024 · 被引用 7 次
