Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing
Blaise Delattre, Hengyu WU, Paul Caillon, Wei Yang Bryan Lim, YANG CAO
摘要
Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continuous and discrete inputs in isolation. This limitation becomes critical in multimodal models, where decisions depend on cross-modal semantics and adversaries can jointly perturb heterogeneous inputs, rendering unimodal certificates insufficient. We introduce a unified randomized smoothing framework for mixed discrete--continuous inputs based on an analytically tractable Neyman--Pearson formulation of the joint worst-case problem. By analyzing the joint likelihood ordering induced by factorized discrete and continuous noise, our approach yields a closed-form, one-dimensional certificate that strictly generalizes both Gaussian (image-only) and discrete (text-only) randomized smoothing. We validate the framework on multimodal safety filtering, providing, to our knowledge, the first model-agnostic Neyman--Pearson certificate for joint discrete-token and continuous-image perturbations in interaction-dependent text--image safety filtering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2023 · 被引用 412 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
相关 Paper
- Higher-Order Certification For Randomized SmoothingJeet Mohapatra, Ching-Yun Ko, Tsui-Wei Weng, Pin-Yu Chen 等NeurIPS 2020 · 被引用 51 次
- Efficient Robustness Certificates for Discrete Data: Sparsity-Aware Randomized Smoothing for Graphs, Images and MoreAleksandar Bojchevski, Johannes Klicpera, Stephan GünnemannICML 2020 · 被引用 95 次
- Localized Randomized Smoothing for Collective Robustness CertificationJan Schuchardt, Tom Wollschläger, Aleksandar Bojchevski, Stephan GünnemannICLR 2023
- Hierarchical Randomized SmoothingYan Scholten, Jan Schuchardt, Aleksandar Bojchevski, Stephan GünnemannNeurIPS 2023 · 被引用 14 次
- Black-Box Certification with Randomized Smoothing: A Functional Optimization Based FrameworkDinghuai Zhang, Mao Ye, Chengyue Gong, Zhanxing Zhu 等NeurIPS 2020 · 被引用 71 次
