How recurrent networks implement contextual processing in sentiment analysis
Niru Maheswaranathan, David Sussillo
Abstract
Neural networks have a remarkable capacity for contextual processing--using recent or nearby inputs to modify processing of current input. For example, in natural language, contextual processing is necessary to correctly interpret negation (e.g. phrases such as "not bad"). However, our ability to understand how networks process context is limited. Here, we propose general methods for reverse engineering recurrent neural networks (RNNs) to identify and elucidate contextual processing. We apply these methods to understand RNNs trained on sentiment classification. This analysis reveals inputs that induce contextual effects, quantifies the strength and timescale of these effects, and identifies sets of these inputs with similar properties. Additionally, we analyze contextual effects related to differential processing of the beginning and end of documents. Using the insights learned from the RNNs we improve baseline Bag-of-Words models with simple extensions that incorporate contextual modification, recovering greater than 90% of the RNN's performance increase over the baseline. This work yields a new understanding of how RNNs process contextual information, and provides tools that should provide similar insight more broadly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Auxiliary Tasks and Exploration Enable ObjectGoal NavigationJoel Ye, Dhruv Batra, Abhishek Das, Erik WijmansICCV 2021 · 137 citations
- Extracting computational mechanisms from neural data using low-rank RNNsAdrian Valente, Jonathan W. Pillow, Srdjan OstojicNeurIPS 2022 · 71 citations
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein et al.NeurIPS 2020 · 56 citations
- Reverse engineering recurrent neural networks with Jacobian switching linear dynamical systemsJimmy T. H. Smith, Scott W. Linderman, David SussilloNeurIPS 2021 · 44 citations
- Understanding How Encoder-Decoder Architectures AttendKyle Aitken, Vinay V. Ramasesh, Yuan Cao, Niru MaheswaranathanNeurIPS 2021 · 32 citations
Related papers
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural NetworkJavier Turek, Shailee Jain, Vy A. Vo, Mihai Capota et al.ICML 2020 · 13 citations
- Word-Level Contextual Sentiment Analysis with InterpretabilityTomoki Ito, Kota Tsubouchi, Hiroki Sakaji, Tatsuo Yamashita et al.AAAI 2020 · 17 citations
- Modeling Hierarchical Structures with Continuous Recursive Neural NetworksJishnu Ray Chowdhury, Cornelia CarageaICML 2021 · 18 citations
- Replicate, Walk, and Stop on Syntax: An Effective Neural Network Model for Aspect-Level Sentiment ClassificationYaowei Zheng, Richong Zhang, Samuel Mensah, Yongyi MaoAAAI 2020 · 48 citations
- TRADER: trace divergence analysis and embedding regulation for debugging recurrent neural networksGuanhong Tao, Shiqing Ma, Yingqi Liu, Qiuling Xu et al.ICSE 2020 · 14 citations
