Efficient Training of Low-Curvature Neural Networks
Suraj Srinivas, Kyle Matoba, Himabindu Lakkaraju, François Fleuret
摘要
Standard deep neural networks often have excess non-linearity, making them susceptible to issues such as low adversarial robustness and gradient instability. Common methods to address these downstream issues, such as adversarial training, are expensive and often sacrifice predictive accuracy. In this work, we address the core issue of excess non-linearity via curvature, and demonstrate low-curvature neural networks (LCNNs) that obtain drastically lower curvature than standard models while exhibiting similar predictive performance. This leads to improved robustness and stable gradients, at a fraction of the cost of standard adversarial training. To achieve this, we decompose overall model curvature in terms of curvatures and slopes of its constituent layers. To enable efficient curvature minimization of constituent layers, we introduce two novel architectural components: first, a non-linearity called centered-softplus that is a stable variant of the softplus non-linearity, and second, a Lipschitz-constrained batch normalization layer. Our experiments show that LCNNs have lower curvature, more stable gradients and increased off-the-shelf adversarial robustness when compared to standard neural networks, all without affecting predictive performance. Our approach is easy to use and can be readily incorporated into existing neural network architectures. Code to implement our method and replicate our experiments is available at https://github.com/kylematoba/lcnn .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Protein Design with Guided Discrete DiffusionNate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner 等NeurIPS 2023 · 被引用 246 次
- Partial Counterfactual Identification of Continuous Outcomes with a Curvature Sensitivity ModelValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelNeurIPS 2023 · 被引用 15 次
- Towards Bridging the Gaps between the Right to Explanation and the Right to be ForgottenSatyapriya Krishna, Jiaqi Ma, Himabindu LakkarajuICML 2023 · 被引用 14 次
- SuperDeepFool: a new fast and accurate minimal adversarial attackAlireza Abdollahpour, Mahed Abroshan, Seyed-Mohsen Moosavi-DezfooliNeurIPS 2024 · 被引用 11 次
- Efficient local linearity regularization to overcome catastrophic overfittingElías Abad-Rocamora, Fanghui Liu, Grigorios Chrysos, Pablo M. Olmos 等ICLR 2024 · 被引用 9 次
它引用的顶会 Paper5
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Orthogonalizing Convolutional Layers with the Cayley TransformAsher Trockman, J. Zico KolterICLR 2021 · 被引用 137 次
- Second-Order Provable Defenses against Adversarial AttacksSahil Singla, Soheil FeiziICML 2020 · 被引用 64 次
- Rethinking the Role of Gradient-based Attribution Methods for Model InterpretabilitySuraj Srinivas, François FleuretICLR 2021 · 被引用 46 次
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji 等USENIX Security 2020
相关 Paper
- Low Curvature Activations Reduce Overfitting in Adversarial TrainingVasu Singla, Sahil Singla, Soheil Feizi, David JacobsICCV 2021 · 被引用 49 次
- Large Norms of CNN Layers Do Not Hurt Adversarial RobustnessYouwei Liang, Dong HuangAAAI 2021 · 被引用 13 次
- Skew Orthogonal ConvolutionsSahil Singla, Soheil FeiziICML 2021 · 被引用 76 次
- Compositional Curvature Bounds for Deep Neural NetworksTaha Entesari, Sina Sharifi, Mahyar FazlyabICML 2024 · 被引用 2 次
- Improved techniques for deterministic l2 robustnessSahil Singla, Soheil FeiziNeurIPS 2022 · 被引用 13 次
