Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction
Guoliang Dong, Jingyi Wang, Jun Sun, Yang Zhang, Xinyu Wang, Ting Dai, Jin Song Dong, Xingen Wang
Abstract
Neural networks are becoming a popular tool for solving many realworld problems such as object recognition and machine translation, thanks to its exceptional performance as an end-to-end solution. However, neural networks are complex black-box models, which hinders humans from interpreting and consequently trusting them in making critical decisions. Towards interpreting neural networks, several approaches have been proposed to extract simple deterministic models from neural networks. The results are not encouraging (e.g., low accuracy and limited scalability), fundamentally due to the limited expressiveness of such simple models. In this work, we propose an approach to extract probabilistic automata for interpreting an important class of neural networks, i.e., recurrent neural networks. Our work distinguishes itself from existing approaches in two important ways. One is that probability is used to compensate for the loss of expressiveness. This is inspired by the observation that human reasoning is often 'probabilistic'. The other is that we adaptively identify the right level of abstraction so that a simple model is extracted in a request-specific way. We conduct experiments on several real-world datasets using state-ofthe-art architectures including GRU and LSTM. The result shows that our approach significantly improves existing approaches in terms of accuracy or scalability. Lastly, we demonstrate the usefulness of the extracted models through detecting adversarial texts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aff08c07-cf9f-4a1b-b3f2-aab23e07cedcCited by top-tier papers2
- Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksXiyue Zhang, Xiaoning Du, Xiaofei Xie, Lei Ma et al.AAAI 2021 · 25 citations
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 6 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
Related papers
- AdaAX: Explaining Recurrent Neural Networks by Learning Automata with Adaptive StatesDat Hong, Alberto Maria Segre, Tong WangKDD 2022 · 3 citations
- Marble: Model-based Robustness Analysis of Stateful Deep Learning SystemsXiaoning Du, Yi Li, Xiaofei Xie, Lei Ma et al.ASE 2020 · 10 citations
- Understanding and Improving Adversarial Robustness of Neural Probabilistic CircuitsWeixin Chen, Han ZhaoNeurIPS 2025 · 1 citation
- Uncertainty Estimation and Calibration with Finite-State Probabilistic RNNsCheng Wang, Carolin Lawrence, Mathias NiepertICLR 2021 · 10 citations
- Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural NetworksChengyue Jiang, Yinggong Zhao, Shanbo Chu, Libin Shen et al.EMNLP 2020 · 23 citations
