Lune

ICML2020Top-tier venue

Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks

David Stutz, Matthias Hein, Bernt Schiele

2020Year
158Citations
39Top-tier citations

Abstract

Adversarial training yields robust models against a specific threat model, e.g., L∞L_\infty adversarial examples. Typically robustness does not generalize to previously unseen threat models, e.g., other LpL_p norms, or larger perturbations. Our confidence-calibrated adversarial training (CCAT) tackles this problem by biasing the model towards low confidence predictions on adversarial examples. By allowing to reject examples with low confidence, robustness generalizes beyond the threat model employed during training. CCAT, trained only on L∞L_\infty adversarial examples, increases robustness against larger L∞L_\infty, L2L_2, L1L_1 and L0L_0 attacks, adversarial frames, distal adversarial examples and corrupted examples and yields better clean accuracy compared to adversarial training. For thorough evaluation we developed novel white- and black-box attacks directly attacking CCAT by maximizing confidence. For each threat model, we use 77 attacks with up to 5050 restarts and 50005000 iterations and report worst-case robust test error, extended to our confidence-thresholded setting, across all attacks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d2e07179-bee0-457e-acfd-3b40ea71eb6c

Cited by top-tier papers39

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines