Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing
Blaise Delattre, Hengyu WU, Paul Caillon, Wei Yang Bryan Lim, YANG CAO
Abstract
Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continuous and discrete inputs in isolation. This limitation becomes critical in multimodal models, where decisions depend on cross-modal semantics and adversaries can jointly perturb heterogeneous inputs, rendering unimodal certificates insufficient. We introduce a unified randomized smoothing framework for mixed discrete--continuous inputs based on an analytically tractable Neyman--Pearson formulation of the joint worst-case problem. By analyzing the joint likelihood ordering induced by factorized discrete and continuous noise, our approach yields a closed-form, one-dimensional certificate that strictly generalizes both Gaussian (image-only) and discrete (text-only) randomized smoothing. We validate the framework on multimodal safety filtering, providing, to our knowledge, the first model-agnostic Neyman--Pearson certificate for joint discrete-token and continuous-image perturbations in interaction-dependent text--image safety filtering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c60d20cd-e1b0-46b5-9d98-6dcad58f1f15Builds on19
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
Related papers
- Higher-Order Certification For Randomized SmoothingJeet Mohapatra, Ching-Yun Ko, Tsui-Wei Weng, Pin-Yu Chen et al.NeurIPS 2020 · 51 citations
- Efficient Robustness Certificates for Discrete Data: Sparsity-Aware Randomized Smoothing for Graphs, Images and MoreAleksandar Bojchevski, Johannes Klicpera, Stephan GünnemannICML 2020 · 95 citations
- Localized Randomized Smoothing for Collective Robustness CertificationJan Schuchardt, Tom Wollschläger, Aleksandar Bojchevski, Stephan GünnemannICLR 2023
- Hierarchical Randomized SmoothingYan Scholten, Jan Schuchardt, Aleksandar Bojchevski, Stephan GünnemannNeurIPS 2023 · 14 citations
- Black-Box Certification with Randomized Smoothing: A Functional Optimization Based FrameworkDinghuai Zhang, Mao Ye, Chengyue Gong, Zhanxing Zhu et al.NeurIPS 2020 · 71 citations
