Provable Gradient Editing of Deep Neural Networks
Zhe Tao, Aditya V. Thakur
Abstract
In explainable AI, DNN gradients are used to interpret the prediction; in safetycritical control systems, gradients could encode safety constraints; in scientificcomputing applications, gradients could encode physical invariants. While recent work on provable editing of DNNs has focused on input-output constraints, the problem of enforcing hard constraints on DNN gradients remains unaddressed. We present ProGrad, the first efficient approach for editing the parameters of a DNN to provably enforce hard constraints on the DNN gradients. Given a DNN N with parameters θ, and a set S of pairs (x x x, Q) of input x x x and corresponding linear gradient constraints Q, ProGrad finds new parameters θ θ θ such that (x x x,Q)∈S ∂ ∂x x x N(x x x; θ θ θ) ∈ Q while minimizing the changes ∥ θ θ θ -θ∥. The key contribution is a novel conditional variable gradient of DNNs, which relaxes the NP-hard provable gradient editing problem to a linear program (LP), enabling ProGrad to use an LP solver to efficiently and effectively enforce the gradient constraints. We experimentally evaluated ProGrad via enforcing (i) hard Grad-CAM constraints on IMAGENET ResNet DNNs; (ii) hard Integrated Gradients constraints on Llama 3 and Qwen 3 LLMs; (iii) hard gradient constraints in training a function-approximation DNN as a proxy for safety constraints in control systems and physical invariants in scientific applications. The results highlight the unique capability of ProGrad in enforcing hard constraints on DNN gradients. class: stingray 1 (a) Original image. Cosine: 0% IoU: 0% 1 (b) Reference Grad-CAM on the original image. misclassified: coral reef 1 (c) Misclassified Gaussian-noise corrupted image. Cos: 34.66% IoU: 4.35% 1 (d) Deviated Grad-CAM on the corrupted image. Cos: 100% IoU: 100%
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e425bb1-631e-4e85-947f-5d7313579852Builds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness VerificationShiqi Wang, Huan Zhang, Kaidi Xu, Xue Lin et al.NeurIPS 2021 · 359 citations
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete VerifiersKaidi Xu, Huan Zhang, Shiqi Wang, Yihan Wang et al.ICLR 2021 · 250 citations
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 186 citations
Related papers
- Provable Editing of Deep Neural Networks using Parametric Linear RelaxationZhe Tao, Aditya V. ThakurNeurIPS 2024 · 5 citations
- SAME: Safety-Aware Model Editing Guided by Safety TransformationJiayi Wang, Shipeng Wang, Ji Wu, Jian SunACL 2026
- GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi Xie, Minghong Fang, Renjie Pi, Neil GongACL 2024
- DeepSaDe: Learning Neural Networks That Guarantee Domain Constraint SatisfactionKshitij Goyal, Sebastijan Dumancic, Hendrik BlockeelAAAI 2024 · 9 citations
- Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsHongpeng Cao, Yanbing Mao, Lui Sha, Marco CaccamoICLR 2024 · 11 citations
