AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
Zhuoqun Huang, Neil G. Marchant, Olga Ohrimenko, Benjamin I. P. Rubinstein
摘要
We consider the problem of certified robustness for sequence classification against edit distance perturbations. Naturally occurring inputs of varying lengths (e.g., sentences in natural language processing tasks) present a challenge to current methods that employ fixed-rate deletion mechanisms and lead to suboptimal performance. To this end, we introduce AdaptDel methods with adaptable deletion rates that dynamically adjust based on input properties. We extend the theoretical framework of randomized smoothing to variable-rate deletion, ensuring sound certification with respect to edit distance. We achieve strong empirical results in natural language tasks, observing up to 30 orders of magnitude improvement to median cardinality of the certified region, over state-of-the-art certifications. potential variations in sequence length, complexity, or domain-specific characteristics. For instance, Huang et al. [18] observed that longer text sequences can typically tolerate higher deletion rates without compromising predictive accuracy. This insight suggests a fixed deletion rate may lead to suboptimal trade-offs between robustness and performance, particularly for inputs that could benefit from more nuanced treatment.
The potential of input-dependent smoothing for discrete sequences remains largely unexplored. Prior work has considered Gaussian smoothing with input-dependent noise for ℓ p certification of realvalued inputs. Other work has taken an ad hoc approach, where Cohen et al.'s certificate for uniform Gaussian noise is corrected iteratively at test time [19,20]. However, this results in a model where each output depends on the chain of previous outputs, which complicates interpretability, may lead to error propagation and increases computational complexity and memory requirements. A more rigorous approach was taken by Súkeník et al. [21], who derived a certificate for input-dependent Gaussian noise. In this paper, we extend input-dependent smoothing to discrete sequences, where a new analysis is required to obtain sound edit distance certificates. We begin by introducing a general framework for deletion smoothing with input-dependent deletion rates, which provides bounds on the smoothed model's confidence scores under perturbations. Building on these bounds, we propose two methods that support efficient edit distance certification: AdaptDel, which uses a length-dependent deletion rate, and AdaptDel+, which further optimizes the deletion rate using input binning and empirical calibration. Our contributions are summarized as follows:
• We develop a theoretical foundation for variable deletion rates, enabling robust smoothing mechanisms that adapt to input length.
• We propose AdaptDel and AdaptDel+, two novel adaptive smoothing techniques, which enable computationally efficient certified robustness while maintaining high accuracies.
• We evaluate our methods on four natural language processing tasks with varying input sizes, demonstrating that AdaptDel and AdaptDel+ achieve stronger robustness on all datasets at the same certified accuracy, with little to no degradation in clean accuracy: for example, we observe up to 30 orders of magnitude improvement to median cardinality of the certified region.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu 等ACL 2020 · 被引用 188 次
相关 Paper
- RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized DeletionZhuoqun Huang, Neil G. Marchant, Keane Lucas, Lujo Bauer 等NeurIPS 2023 · 被引用 24 次
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
- Higher-Order Certification For Randomized SmoothingJeet Mohapatra, Ching-Yun Ko, Tsui-Wei Weng, Pin-Yu Chen 等NeurIPS 2020 · 被引用 51 次
- Center Smoothing: Certified Robustness for Networks with Structured OutputsAounon Kumar, Tom GoldsteinNeurIPS 2021 · 被引用 23 次
- Robustness Certificates for Sparse Adversarial Attacks by Randomized AblationAlexander Levine, Soheil FeiziAAAI 2020 · 被引用 114 次
