Benchmarking Deep Learning Interpretability in Time Series Predictions
Aya Abdelsalam Ismail, Mohamed K. Gunady, Héctor Corrada Bravo, Soheil Feizi
Abstract
Saliency methods are used extensively to highlight the importance of input features in model predictions. These methods are mostly used in vision and language tasks, and their applications to time series data is relatively unexplored. In this paper, we set out to extensively compare the performance of various saliency-based interpretability methods across diverse neural architectures, including Recurrent Neural Network, Temporal Convolutional Networks, and Transformers in a new benchmark † of synthetic time series data. We propose and report multiple metrics to empirically evaluate the performance of saliency methods for detecting feature importance over time using both precision (i.e., whether identified features contain meaningful signals) and recall (i.e., the number of features with signal identified as important). Through several experiments, we show that (i) in general, network architectures and saliency methods fail to reliably and accurately identify feature importance over time in time series data, (ii) this failure is mainly due to the conflation of time and feature domains, and (iii) the quality of saliency maps can be improved substantially by using our proposed two-step temporal saliency rescaling (TSR) approach that first calculates the importance of each time step before calculating the importance of each feature at a time step.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b083c3cf-4f3b-430c-9a82-2e8bec21b465Cited by top-tier papers27
- Generative Time Series Forecasting with Diffusion, Denoise, and DisentanglementYan Li, Xinjiang Lu, Yaqing Wang, Dejing DouNeurIPS 2022 · 203 citations
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 167 citations
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 121 citations
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 115 citations
Related papers
- Temporal Dependencies in Feature Importance for Time Series PredictionKin Kwan Leung, Clayton Rooke, Jonathan Smith, Saba Zuberi et al.ICLR 2023 · 7 citations
- Learning Perturbations to Explain Time Series PredictionsJoseph EnguehardICML 2023 · 29 citations
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
- Time series saliency maps: Explaining models across multiple domainsChristodoulos Kechris, Jonathan Dan, David AtienzaICML 2026 · 6 citations
- Feature Importance Explanations for Temporal Black-Box ModelsAkshay Sood, Mark CravenAAAI 2022 · 24 citations
