Provable Gradient Editing of Deep Neural Networks
Zhe Tao, Aditya V. Thakur
摘要
In explainable AI, DNN gradients are used to interpret the prediction; in safetycritical control systems, gradients could encode safety constraints; in scientificcomputing applications, gradients could encode physical invariants. While recent work on provable editing of DNNs has focused on input-output constraints, the problem of enforcing hard constraints on DNN gradients remains unaddressed. We present ProGrad, the first efficient approach for editing the parameters of a DNN to provably enforce hard constraints on the DNN gradients. Given a DNN N with parameters θ, and a set S of pairs (x x x, Q) of input x x x and corresponding linear gradient constraints Q, ProGrad finds new parameters θ θ θ such that (x x x,Q)∈S ∂ ∂x x x N(x x x; θ θ θ) ∈ Q while minimizing the changes ∥ θ θ θ -θ∥. The key contribution is a novel conditional variable gradient of DNNs, which relaxes the NP-hard provable gradient editing problem to a linear program (LP), enabling ProGrad to use an LP solver to efficiently and effectively enforce the gradient constraints. We experimentally evaluated ProGrad via enforcing (i) hard Grad-CAM constraints on IMAGENET ResNet DNNs; (ii) hard Integrated Gradients constraints on Llama 3 and Qwen 3 LLMs; (iii) hard gradient constraints in training a function-approximation DNN as a proxy for safety constraints in control systems and physical invariants in scientific applications. The results highlight the unique capability of ProGrad in enforcing hard constraints on DNN gradients. class: stingray 1 (a) Original image. Cosine: 0% IoU: 0% 1 (b) Reference Grad-CAM on the original image. misclassified: coral reef 1 (c) Misclassified Gaussian-noise corrupted image. Cos: 34.66% IoU: 4.35% 1 (d) Deviated Grad-CAM on the corrupted image. Cos: 100% IoU: 100%
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness VerificationShiqi Wang, Huan Zhang, Kaidi Xu, Xue Lin 等NeurIPS 2021 · 被引用 359 次
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete VerifiersKaidi Xu, Huan Zhang, Shiqi Wang, Yihan Wang 等ICLR 2021 · 被引用 250 次
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 被引用 186 次
相关 Paper
- Provable Editing of Deep Neural Networks using Parametric Linear RelaxationZhe Tao, Aditya V. ThakurNeurIPS 2024 · 被引用 5 次
- SAME: Safety-Aware Model Editing Guided by Safety TransformationJiayi Wang, Shipeng Wang, Ji Wu, Jian SunACL 2026
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi Xie, Minghong Fang, Renjie Pi, Neil GongACL 2024
- DeepSaDe: Learning Neural Networks That Guarantee Domain Constraint SatisfactionKshitij Goyal, Sebastijan Dumancic, Hendrik BlockeelAAAI 2024 · 被引用 9 次
- Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsHongpeng Cao, Yanbing Mao, Lui Sha, Marco CaccamoICLR 2024 · 被引用 11 次
