AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
Zhuoqun Huang, Neil G. Marchant, Olga Ohrimenko, Benjamin I. P. Rubinstein
Abstract
We consider the problem of certified robustness for sequence classification against edit distance perturbations. Naturally occurring inputs of varying lengths (e.g., sentences in natural language processing tasks) present a challenge to current methods that employ fixed-rate deletion mechanisms and lead to suboptimal performance. To this end, we introduce AdaptDel methods with adaptable deletion rates that dynamically adjust based on input properties. We extend the theoretical framework of randomized smoothing to variable-rate deletion, ensuring sound certification with respect to edit distance. We achieve strong empirical results in natural language tasks, observing up to 30 orders of magnitude improvement to median cardinality of the certified region, over state-of-the-art certifications. potential variations in sequence length, complexity, or domain-specific characteristics. For instance, Huang et al. [18] observed that longer text sequences can typically tolerate higher deletion rates without compromising predictive accuracy. This insight suggests a fixed deletion rate may lead to suboptimal trade-offs between robustness and performance, particularly for inputs that could benefit from more nuanced treatment.
The potential of input-dependent smoothing for discrete sequences remains largely unexplored. Prior work has considered Gaussian smoothing with input-dependent noise for ℓ p certification of realvalued inputs. Other work has taken an ad hoc approach, where Cohen et al.'s certificate for uniform Gaussian noise is corrected iteratively at test time [19,20]. However, this results in a model where each output depends on the chain of previous outputs, which complicates interpretability, may lead to error propagation and increases computational complexity and memory requirements. A more rigorous approach was taken by Súkeník et al. [21], who derived a certificate for input-dependent Gaussian noise. In this paper, we extend input-dependent smoothing to discrete sequences, where a new analysis is required to obtain sound edit distance certificates. We begin by introducing a general framework for deletion smoothing with input-dependent deletion rates, which provides bounds on the smoothed model's confidence scores under perturbations. Building on these bounds, we propose two methods that support efficient edit distance certification: AdaptDel, which uses a length-dependent deletion rate, and AdaptDel+, which further optimizes the deletion rate using input binning and empirical calibration. Our contributions are summarized as follows:
• We develop a theoretical foundation for variable deletion rates, enabling robust smoothing mechanisms that adapt to input length.
• We propose AdaptDel and AdaptDel+, two novel adaptive smoothing techniques, which enable computationally efficient certified robustness while maintaining high accuracies.
• We evaluate our methods on four natural language processing tasks with varying input sizes, demonstrating that AdaptDel and AdaptDel+ achieve stronger robustness on all datasets at the same certified accuracy, with little to no degradation in clean accuracy: for example, we observe up to 30 orders of magnitude improvement to median cardinality of the certified region.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
Related papers
- RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized DeletionZhuoqun Huang, Neil G. Marchant, Keane Lucas, Lujo Bauer et al.NeurIPS 2023 · 24 citations
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang et al.S&P 2024 · 41 citations
- Higher-Order Certification For Randomized SmoothingJeet Mohapatra, Ching-Yun Ko, Tsui-Wei Weng, Pin-Yu Chen et al.NeurIPS 2020 · 51 citations
- Center Smoothing: Certified Robustness for Networks with Structured OutputsAounon Kumar, Tom GoldsteinNeurIPS 2021 · 23 citations
- Robustness Certificates for Sparse Adversarial Attacks by Randomized AblationAlexander Levine, Soheil FeiziAAAI 2020 · 114 citations
