How Does Selective Mechanism Improve Self-Attention Networks?
Xinwei Geng, Longyue Wang, Xing Wang, Bing Qin, Ting Liu, Zhaopeng Tu
摘要
Self-attention networks (SANs) with selective mechanism has produced substantial improvements in various NLP tasks by concentrating on a subset of input words. However, the underlying reasons for their strong performance have not been well explained. In this paper, we bridge the gap by assessing the strengths of selective SANs (SSANs), which are implemented with a flexible and universal Gumbel-Softmax. Experimental results on several representative NLP tasks, including natural language inference, semantic role labelling, and machine translation, show that SSANs consistently outperform the standard SANs. Through well-designed probing experiments, we empirically validate that the improvement of SSANs can be attributed in part to mitigating two commonly-cited weaknesses of SANs: word order encoding and structure modeling. Specifically, the selective mechanism improves SANs by paying more attention to content words that contribute to the meaning of the sentence. The code and data are released at https://github.com/xwgeng/SSAN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Improving Visual Prompt Tuning for Self-supervised Vision TransformersSeungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee 等ICML 2023 · 被引用 74 次
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song 等NeurIPS 2024 · 被引用 41 次
- Towards Trustworthy Explanation: On Causal RationalizationWenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai 等ICML 2023 · 被引用 25 次
- Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence ModelingHongyu Gong, Yun Tang, Juan Miguel Pino, Xian LiNeurIPS 2021 · 被引用 15 次
- Learning to Select Prototypical Parts for Interpretable Sequential Data ModelingYifei Zhang, Neng Gao, Cunqing MaAAAI 2023 · 被引用 10 次
相关 Paper
- Transformer Uncertainty Estimation with Hierarchical Stochastic AttentionJiahuan Pei, Cheng Wang, György SzarvasAAAI 2022 · 被引用 33 次
- Semantics-Aware Inferential Network for Natural Language UnderstandingShuailiang Zhang, Hai Zhao, Junru Zhou, Xi Zhou 等AAAI 2021 · 被引用 4 次
- Interpreting Positional Information in Perspective of Word OrderXilong Zhang, Ruochen Liu, Jin Liu, Xuefeng LiangACL 2023
- InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature InterpretationJacob Yoke Hong Si, Wendy Yusi Cheng, Michael Cooper, Rahul G. KrishnanICML 2024 · 被引用 15 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
