Learning Multiscale Transformer Models for Sequence Generation
Bei Li, Tong Zheng, Yi Jing, Chengbo Jiao, Tong Xiao, Jingbo Zhu
Abstract
Multiscale feature hierarchies have been witnessed the success in the computer vision area. This further motivates researchers to design multiscale Transformer for natural language processing, mostly based on the self-attention mechanism. For example, restricting the receptive field across heads or extracting local fine-grained features via convolutions. However, most of existing works directly modeled local features but ignored the word-boundary information. This results in redundant and ambiguous attention distributions, which lacks of interpretability. In this work, we define those scales in different linguistic units, including sub-words, words and phrases. We built a multiscale Transformer model by establishing relationships among scales based on word-boundary information and phrase-level prior knowledge. The proposed Universal MultiScale Transformer, namely Umst, was evaluated on two sequence generation tasks. Notably, it yielded consistent performance gains over the strong baseline on several test sets without sacrificing the efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 320eea1f-5da3-4e51-adb3-836af8c97e73Cited by top-tier papers3
- ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence GenerationChenglong Wang, Hang Zhou, Yimin Hu, Yifu Huo et al.AAAI 2024 · 15 citations
- EIT: Enhanced Interactive TransformerTong Zheng, Bei Li, Huiwen Bao, Tong Xiao et al.ACL 2024
- Sparse-Scale Transformer with Bidirectional Awareness for Time Series ForecastingYing Liu, Bo Liu, Sheng Huang, Gang Luo et al.AAAI 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
Related papers
- Multi-Scale Self-Attention for Text ClassificationQipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue et al.AAAI 2020 · 69 citations
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio et al.ACL 2023
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
- Learning Source Phrase Representations for Neural Machine TranslationHongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu et al.ACL 2020 · 18 citations
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
