Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust Classifiers
Jonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso, Christian Gagné
Abstract
Despite extensive research on adversarial training strategies to improve robustness, the decisions of even the most robust deep learning models can still be quite sensitive to imperceptible perturbations, creating serious risks when deploying them for high-stakes real-world applications. While detecting such cases may be critical, evaluating a model's vulnerability at a per-instance level using adversarial attacks is computationally too intensive and unsuitable for real-time deployment scenarios. The input space margin is the exact score to detect non-robust samples and is intractable for deep neural networks. This paper introduces the concept of margin consistency -- a property that links the input space margins and the logit margins in robust models -- for efficient detection of vulnerable samples. First, we establish that margin consistency is a necessary and sufficient condition to use a model's logit margin as a score for identifying non-robust samples. Next, through comprehensive empirical analysis of various robustly trained models on CIFAR10 and CIFAR100 datasets, we show that they indicate high margin consistency with a strong correlation between their input space margins and the logit margins. Then, we show that we can effectively and confidently use the logit margin to detect brittle decisions with such models. Finally, we address cases where the model is not sufficiently margin-consistent by learning a pseudo-margin from the feature representation. Our findings highlight the potential of leveraging deep representations to assess adversarial vulnerability in deployment scenarios efficiently.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0496e359-e44c-414b-a358-07051b853377Cited by top-tier papers2
- FLaG: Fine-Grained Latent Grouping for Hallucination DetectionWentao Ye, Liyao Li, Zhiqing Xiao, Muzhi Zhu et al.KDD 2026 · 1 citation
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Epsilon-SchedulingJonas Ngnawé, Maxime Heuillet, Sabyasachi Sahoo, Yann Pequignot et al.ICLR 2026
Builds on25
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual AttacksYong Xie, Weijie Zheng, Hanxun Huang, Guangnan Ye et al.CVPR 2025
- Exploring and Exploiting Decision Boundary Dynamics for Adversarial RobustnessYuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein et al.ICLR 2023 · 10 citations
- Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz RegularizationMahyar Fazlyab, Taha Entesari, Aniket Roy, Rama ChellappaNeurIPS 2023 · 26 citations
- Learnable Boundary Guided Adversarial TrainingJiequan Cui, Shu Liu, Liwei Wang, Jiaya JiaICCV 2021 · 152 citations
- Vulnerable Data-Aware Adversarial TrainingYuqi Feng, Jiahao Fan, Yanan SunNeurIPS 2025 · 2 citations
