Adaptive wavelet distillation from neural networks through interpretations
Wooseok Ha, Chandan Singh, François Lanusse, Srigokul Upadhyayula, Bin Yu
Abstract
Recent deep-learning models have achieved impressive prediction performance, but often sacrifice interpretability and computational efficiency. Interpretability is crucial in many disciplines, such as science and medicine, where models must be carefully vetted or where interpretation is the goal itself. Moreover, interpretable models are concise and often yield computational efficiency. Here, we propose adaptive wavelet distillation (AWD), a method which aims to distill information from a trained neural network into a wavelet transform. Specifically, AWD penalizes feature attributions of a neural network in the wavelet domain to learn an effective multi-resolution wavelet transform. The resulting model is highly predictive, concise, computationally efficient, and has properties (such as a multi-scale structure) which make it easy to interpret. In close collaboration with domain experts, we showcase how AWD addresses challenges in two real-world settings: cosmological parameter inference and molecular-partner prediction. In both cases, AWD yields a scientifically interpretable and concise model which gives predictive performance better than state-of-the-art neural networks. Moreover, AWD identifies predictive features that are scientifically meaningful in the context of respective domains. All code and models are released in a full-fledged package available on Github (https://github.com/Yu-Group/adaptive-wavelets).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3deaf70d-ee61-470b-a91a-662e7f3c5299Cited by top-tier papers5
- Towards Faithful XAI Evaluation via Generalization-Limited Backdoor WatermarkMengxi Ya, Yiming Li, Tao Dai, Bin Wang et al.ICLR 2024 · 19 citations
- Interpretable Next-token Prediction via the Generalized Induction HeadEunji Kim, Sriya Mantena, Weiwei Yang, Chandan Singh et al.NeurIPS 2025 · 3 citations
- Debiasing Trace Guidance: Top-Down Trace Distillation and Bottom-up Velocity Alignment for Unsupervised Anomaly DetectionXingjian Wang, Li Chai, Jiming ChenICCV 2025 · 2 citations
- Efficient Multi-Scale Network with Learnable Discrete Wavelet Transform for Blind Motion DeblurringXin Gao, Tianheng Qiu, Xinyu Zhang, Hanlin Bai et al.CVPR 2024
- One Wave To Explain Them All: A Unifying Perspective On Feature AttributionGabriel Kasmi, Amandine Brunetto, Thomas Fel, Jayneel ParekhICML 2025
Builds on1
Related papers
- Task-Driven Causal Feature Distillation: Towards Trustworthy Risk PredictionZhixuan Chu, Mengxuan Hu, Qing Cui, Longfei Li et al.AAAI 2024 · 13 citations
- Wavelet Knowledge Distillation: Towards Efficient Image-to-Image TranslationLinfeng Zhang, Xin Chen, Xiaobing Tu, Pengfei Wan et al.CVPR 2022 · 105 citations
- A Knowledge Distillation-Based Approach to Enhance Transparency of Classifier ModelsYuchen Jiang, Xinyuan Zhao, Yihang Wu, Ahmad ChaddadAAAI 2025 · 5 citations
- FEAT-KD: Learning Concise Representations for Single and Multi-Target Regression via TabNet Knowledge DistillationKei Sen Fong, Mehul MotaniICML 2025
- Knowledge Distillation via Constrained Variational InferenceArdavan Saeedi, Yuria Utsumi, Li Sun, Kayhan Batmanghelich et al.AAAI 2022 · 4 citations
