Segatron: Segment-Aware Transformer for Language Modeling and Understanding
He Bai, Peng Shi, Jimmy Lin, Yuqing Xie, Luchen Tan, Kun Xiong, Wen Gao, Ming Li
Abstract
Transformers are powerful for sequence modeling. Nearly all state-of-the-art language models and pre-trained language models are based on the Transformer architecture. However, it distinguishes sequential tokens only with the token position index. We hypothesize that better contextual representations can be generated from the Transformer with richer positional information. To verify this, we propose a segment-aware Transformer (Segatron), by replacing the original token position encoding with a combined position encoding of paragraph, sentence, and token. We first introduce the segment-aware mechanism to Transformer-XL, which is a popular Transformer-based language model with memory extension and relative position encoding. We find that our method can further improve the Transformer-XL base model and large model, achieving 17.1 perplexity on the WikiText-103 dataset. We further investigate the pre-training masked language modeling task with Segatron. Experimental results show that BERT pre-trained with Segatron (SegaBERT) can outperform BERT with vanilla Transformer on various NLP tasks, and outperforms RoBERTa on zero-shot sentence representation learning. Our code is available on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d3b6947-2f49-4171-95eb-6f13cffe63edCited by top-tier papers8
- A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingHe Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma et al.ICML 2022 · 64 citations
- GNN-LM: Language Modeling based on Global Contexts via GNNYuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun et al.ICLR 2022 · 46 citations
- Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length ExtrapolationZhenyu He, Guhao Feng, Shengjie Luo, Kai Yang et al.ICML 2024 · 25 citations
- Efficient Transformers with Dynamic Token PoolingPiotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria PontiACL 2023 · 14 citations
- Unsupervised Dependency Graph NetworkYikang Shen, Shawn Tan, Alessandro Sordoni, Peng Li et al.ACL 2022 · 8 citations
Builds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu et al.ICLR 2020 · 297 citations
Related papers
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- Investigating Efficiently Extending Transformers for Long Input SummarizationJason Phang, Yao Zhao, Peter J. LiuEMNLP 2023 · 30 citations
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li et al.AAAI 2020 · 396 citations
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang et al.AAAI 2024 · 3 citations
