Learning with Explanation Constraints
Rattana Pukdee, Dylan Sam, J. Zico Kolter, Maria-Florina Balcan, Pradeep Ravikumar
Abstract
As larger deep learning models are hard to interpret, there has been a recent focus on generating explanations of these black-box models. In contrast, we may have apriori explanations of how models should behave. In this paper, we formalize this notion as learning from explanation constraints and provide a learning theoretic framework to analyze how such explanations can improve the learning of our models. One may naturally ask, "When would these explanations be helpful?" Our first key contribution addresses this question via a class of models that satisfies these explanation constraints in expectation over new data. We provide a characterization of the benefits of these models (in terms of the reduction of their Rademacher complexities) for a canonical class of explanations given by gradient information in the settings of both linear models and two layer neural networks. In addition, we provide an algorithmic solution for our framework, via a variational approximation that achieves better performance and satisfies these constraints more frequently, when compared to simpler augmented Lagrangian methods to incorporate these explanations. We demonstrate the benefits of our approach over a large array of synthetic and real-world experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f90d777e-faf2-46fb-a944-e23af4305e7cCited by top-tier papers5
- Expanding the Capabilities of Reinforcement Learning via Text FeedbackYuda Song, Lili Chen, Fahim Tajwar, REMI MUNOS et al.ICML 2026 · 41 citations
- Saliency strikes back: How filtering out high frequencies improves white-box explanationsSabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol et al.ICML 2024 · 4 citations
- Learning from Interval TargetsRattana Pukdee, Ziqi Ke, Chirag GuptaNeurIPS 2025 · 1 citation
- Understanding the Impact of Introducing Constraints at Inference Time on Generalization ErrorMasaaki Nishino, Kengo Nakamura, Norihito YasudaICML 2024 · 1 citation
- Learning from weak labelers as constraintsVishwajeet Agrawal, Rattana Pukdee, Maria-Florina Balcan, Pradeep Kumar RavikumarICLR 2025
Builds on12
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 261 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 121 citations
Related papers
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 46 citations
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- CF-OPT: Counterfactual Explanations for Structured PredictionGermain Vivier-Ardisson, Alexandre Forel, Axel Parmentier, Thibaut VidalICML 2024 · 3 citations
