Multi-Scale Self-Attention for Text Classification
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue, Zheng Zhang
Abstract
In this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e947526d-710b-47f6-b0c2-0e8c6181eaf0Cited by top-tier papers3
- Molformer: Motif-Based Transformer on 3D Heterogeneous Molecular GraphsFang Wu, Dragomir Radev, Stan Z. LiAAAI 2023 · 96 citations
- Faster Depth-Adaptive TransformersYijin Liu, Fandong Meng, Jie Zhou, Yufeng Chen et al.AAAI 2021 · 44 citations
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao et al.ICML 2022 · 15 citations
Related papers
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
- CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale AttentionWenxiao Wang, Lu Yao, Long Chen, Binbin Lin et al.ICLR 2022 · 367 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu et al.ICLR 2024 · 197 citations
- Scaleformer: Iterative Multi-scale Refining Transformers for Time Series ForecastingMohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan SylvainICLR 2023 · 36 citations
