Understanding How Encoder-Decoder Architectures Attend
Kyle Aitken, Vinay V. Ramasesh, Yuan Cao, Niru Maheswaranathan
Abstract
Encoder-decoder networks with attention have proven to be a powerful way to solve many sequence-to-sequence tasks. In these networks, attention aligns encoder and decoder states and is often used for visualizing network behavior. However, the mechanisms used by networks to generate appropriate attention matrices are still mysterious. Moreover, how these mechanisms vary depending on the particular architecture used for the encoder and decoder (recurrent, feed-forward, etc.) are also not well understood. In this work, we investigate how encoder-decoder networks solve different sequence-to-sequence tasks. We introduce a way of decomposing hidden states over a sequence into temporal (independent of input) and input-driven (independent of sequence position) components. This reveals how attention matrices are formed: depending on the task requirements, networks rely more heavily on either the temporal or input-driven components. These findings hold across both recurrent and feed-forward architectures despite their differences in forming the temporal components. Overall, our results provide new insight into the inner workings of attention-based encoder-decoder networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14d27e7b-6773-47ec-bb3b-da650bd4aa0dBuilds on4
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- How recurrent networks implement contextual processing in sentiment analysisNiru Maheswaranathan, David SussilloICML 2020 · 25 citations
- The geometry of integration in text classification RNNsKyle Aitken, Vinay Venkatesh Ramasesh, Ankush Garg, Yuan Cao et al.ICLR 2021 · 4 citations
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
Related papers
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 8 citations
- Location Attention for Extrapolation to Longer SequencesYann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia BruniACL 2020 · 2 citations
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 47 citations
- WHEN: A Wavelet-DTW Hybrid Attention Network for Heterogeneous Time Series AnalysisJingyuan Wang, Chen Yang, Xiaohan Jiang, Junjie WuKDD 2023 · 28 citations
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong et al.AAAI 2021 · 2 citations
