Understanding How Encoder-Decoder Architectures Attend
Kyle Aitken, Vinay V. Ramasesh, Yuan Cao, Niru Maheswaranathan
摘要
Encoder-decoder networks with attention have proven to be a powerful way to solve many sequence-to-sequence tasks. In these networks, attention aligns encoder and decoder states and is often used for visualizing network behavior. However, the mechanisms used by networks to generate appropriate attention matrices are still mysterious. Moreover, how these mechanisms vary depending on the particular architecture used for the encoder and decoder (recurrent, feed-forward, etc.) are also not well understood. In this work, we investigate how encoder-decoder networks solve different sequence-to-sequence tasks. We introduce a way of decomposing hidden states over a sequence into temporal (independent of input) and input-driven (independent of sequence position) components. This reveals how attention matrices are formed: depending on the task requirements, networks rely more heavily on either the temporal or input-driven components. These findings hold across both recurrent and feed-forward architectures despite their differences in forming the temporal components. Overall, our results provide new insight into the inner workings of attention-based encoder-decoder networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield 等ACL 2020 · 被引用 132 次
- How recurrent networks implement contextual processing in sentiment analysisNiru Maheswaranathan, David SussilloICML 2020 · 被引用 25 次
- The geometry of integration in text classification RNNsKyle Aitken, Vinay Venkatesh Ramasesh, Ankush Garg, Yuan Cao 等ICLR 2021 · 被引用 4 次
- Transformer Interpretability Beyond Attention VisualizationHila Chefer, Shir Gur, Lior WolfCVPR 2021
相关 Paper
- On the approximation properties of recurrent encoder-decoder architecturesZhong Li, Haotian Jiang, Qianxiao LiICLR 2022 · 被引用 8 次
- Location Attention for Extrapolation to Longer SequencesYann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia BruniACL 2020 · 被引用 2 次
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 被引用 47 次
- WHEN: A Wavelet-DTW Hybrid Attention Network for Heterogeneous Time Series AnalysisJingyuan Wang, Chen Yang, Xiaohan Jiang, Junjie WuKDD 2023 · 被引用 28 次
- Multi-Task Recurrent Modular NetworksDongkuan Xu, Wei Cheng, Xin Dong, Bo Zong 等AAAI 2021 · 被引用 2 次
