Multi-Scale Self-Attention for Text Classification
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue, Zheng Zhang
2020年份
69被引次数
3顶会引用
摘要
In this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Molformer: Motif-Based Transformer on 3D Heterogeneous Molecular GraphsFang Wu, Dragomir Radev, Stan Z. LiAAAI 2023 · 被引用 96 次
- Faster Depth-Adaptive TransformersYijin Liu, Fandong Meng, Jie Zhou, Yufeng Chen 等AAAI 2021 · 被引用 44 次
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao 等ICML 2022 · 被引用 15 次
相关 Paper
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
- CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale AttentionWenxiao Wang, Lu Yao, Long Chen, Binbin Lin 等ICLR 2022 · 被引用 367 次
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu 等ICLR 2024 · 被引用 197 次
- Scaleformer: Iterative Multi-scale Refining Transformers for Time Series ForecastingMohammad Amin Shabani, Amir H. Abdi, Lili Meng, Tristan SylvainICLR 2023 · 被引用 36 次
