How recurrent networks implement contextual processing in sentiment analysis
Niru Maheswaranathan, David Sussillo
摘要
Neural networks have a remarkable capacity for contextual processing--using recent or nearby inputs to modify processing of current input. For example, in natural language, contextual processing is necessary to correctly interpret negation (e.g. phrases such as "not bad"). However, our ability to understand how networks process context is limited. Here, we propose general methods for reverse engineering recurrent neural networks (RNNs) to identify and elucidate contextual processing. We apply these methods to understand RNNs trained on sentiment classification. This analysis reveals inputs that induce contextual effects, quantifies the strength and timescale of these effects, and identifies sets of these inputs with similar properties. Additionally, we analyze contextual effects related to differential processing of the beginning and end of documents. Using the insights learned from the RNNs we improve baseline Bag-of-Words models with simple extensions that incorporate contextual modification, recovering greater than 90% of the RNN's performance increase over the baseline. This work yields a new understanding of how RNNs process contextual information, and provides tools that should provide similar insight more broadly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Auxiliary Tasks and Exploration Enable ObjectGoal NavigationJoel Ye, Dhruv Batra, Abhishek Das, Erik WijmansICCV 2021 · 被引用 137 次
- Extracting computational mechanisms from neural data using low-rank RNNsAdrian Valente, Jonathan W. Pillow, Srdjan OstojicNeurIPS 2022 · 被引用 71 次
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein 等NeurIPS 2020 · 被引用 56 次
- Reverse engineering recurrent neural networks with Jacobian switching linear dynamical systemsJimmy T. H. Smith, Scott W. Linderman, David SussilloNeurIPS 2021 · 被引用 44 次
- Understanding How Encoder-Decoder Architectures AttendKyle Aitken, Vinay V. Ramasesh, Yuan Cao, Niru MaheswaranathanNeurIPS 2021 · 被引用 32 次
相关 Paper
- Approximating Stacked and Bidirectional Recurrent Architectures with the Delayed Recurrent Neural NetworkJavier Turek, Shailee Jain, Vy A. Vo, Mihai Capota 等ICML 2020 · 被引用 13 次
- Word-Level Contextual Sentiment Analysis with InterpretabilityTomoki Ito, Kota Tsubouchi, Hiroki Sakaji, Tatsuo Yamashita 等AAAI 2020 · 被引用 17 次
- Modeling Hierarchical Structures with Continuous Recursive Neural NetworksJishnu Ray Chowdhury, Cornelia CarageaICML 2021 · 被引用 18 次
- Replicate, Walk, and Stop on Syntax: An Effective Neural Network Model for Aspect-Level Sentiment ClassificationYaowei Zheng, Richong Zhang, Samuel Mensah, Yongyi MaoAAAI 2020 · 被引用 48 次
- TRADER: trace divergence analysis and embedding regulation for debugging recurrent neural networksGuanhong Tao, Shiqing Ma, Yingqi Liu, Qiuling Xu 等ICSE 2020 · 被引用 14 次
