Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction
Guoliang Dong, Jingyi Wang, Jun Sun, Yang Zhang, Xinyu Wang, Ting Dai, Jin Song Dong, Xingen Wang
摘要
Neural networks are becoming a popular tool for solving many realworld problems such as object recognition and machine translation, thanks to its exceptional performance as an end-to-end solution. However, neural networks are complex black-box models, which hinders humans from interpreting and consequently trusting them in making critical decisions. Towards interpreting neural networks, several approaches have been proposed to extract simple deterministic models from neural networks. The results are not encouraging (e.g., low accuracy and limited scalability), fundamentally due to the limited expressiveness of such simple models. In this work, we propose an approach to extract probabilistic automata for interpreting an important class of neural networks, i.e., recurrent neural networks. Our work distinguishes itself from existing approaches in two important ways. One is that probability is used to compensate for the loss of expressiveness. This is inspired by the observation that human reasoning is often 'probabilistic'. The other is that we adaptively identify the right level of abstraction so that a simple model is extracted in a request-specific way. We conduct experiments on several real-world datasets using state-ofthe-art architectures including GRU and LSTM. The result shows that our approach significantly improves existing approaches in terms of accuracy or scalability. Lastly, we demonstrate the usefulness of the extracted models through detecting adversarial texts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Decision-Guided Weighted Automata Extraction from Recurrent Neural NetworksXiyue Zhang, Xiaoning Du, Xiaofei Xie, Lei Ma 等AAAI 2021 · 被引用 25 次
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 被引用 6 次
它引用的顶会 Paper3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov 等S&P 2018 · 被引用 987 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
相关 Paper
- AdaAX: Explaining Recurrent Neural Networks by Learning Automata with Adaptive StatesDat Hong, Alberto Maria Segre, Tong WangKDD 2022 · 被引用 3 次
- Marble: Model-based Robustness Analysis of Stateful Deep Learning SystemsXiaoning Du, Yi Li, Xiaofei Xie, Lei Ma 等ASE 2020 · 被引用 10 次
- Understanding and Improving Adversarial Robustness of Neural Probabilistic CircuitsWeixin Chen, Han ZhaoNeurIPS 2025 · 被引用 1 次
- Uncertainty Estimation and Calibration with Finite-State Probabilistic RNNsCheng Wang, Carolin Lawrence, Mathias NiepertICLR 2021 · 被引用 10 次
- Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural NetworksChengyue Jiang, Yinggong Zhao, Shanbo Chu, Libin Shen 等EMNLP 2020 · 被引用 23 次
