Multiple Positional Self-Attention Network for Text Classification
Biyun Dai, Jinlong Li, Ruoyi Xu
Abstract
Self-attention mechanisms have recently caused many concerns on Natural Language Processing (NLP) tasks. Relative positional information is important to self-attention mechanisms. We propose Faraway Mask focusing on the (2m + 1)-gram words and Scaled-Distance Mask putting the logarithmic distance punishment to avoid and weaken the self-attention of distant words respectively. To exploit different masks, we present Positional Self-Attention Layer for generating different Masked-Self-Attentions and a following Position-Fusion Layer in which fused positional information multiplies the Masked-Self-Attentions for generating sentence embeddings. To evaluate our sentence embeddings approach Multiple Positional Self-Attention Network (MPSAN), we perform the comparison experiments on sentiment analysis, semantic relatedness and sentence classification tasks. The result shows that our MPSAN outperforms state-of-the-art methods on five datasets and the test accuracy is improved by 0.81%, 0.6% on SST, CR datasets, respectively. In addition, we reduce training parameters and improve the time efficiency of MPSAN by lowering the dimension number of self-attention and simplifying fusion mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f6925cd-eb9c-4bf4-86a4-3a76d85f179aCited by top-tier papers3
- ClusterFormer: Neural Clustering Attention for Efficient and Effective TransformerNingning Wang, Guobing Gan, Peng Zhang, Shuai Zhang et al.ACL 2022 · 24 citations
- Continuous Self-Attention Models with Neural ODE NetworksJing Zhang, Peng Zhang, Baiwen Kong, Junqiu Wei et al.AAAI 2021 · 24 citations
- On Scalar Embedding of Relative Positions in Attention ModelsJunshuang Wu, Richong Zhang, Yongyi Mao, Junfan ChenAAAI 2021 · 4 citations
Related papers
- Modelling Context and Syntactical Features for Aspect-based Sentiment AnalysisMinh-Hieu Phan, Philip O. OgunbonaACL 2020 · 190 citations
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang et al.ICLR 2021 · 129 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- Discourse Self-Attention for Discourse Element Identification in Argumentative Student EssaysWei Song, Ziyao Song, Ruiji Fu, Lizhen Liu et al.EMNLP 2020 · 15 citations
- Self-Attention Enhanced Selective Gate with Entity-Aware Embedding for Distantly Supervised Relation ExtractionYang Li, Guodong Long, Tao Shen, Tianyi Zhou et al.AAAI 2020 · 88 citations
