Positional Artefacts Propagate Through Masked Language Model Embeddings
Ziyang Luo, Artur Kulmizev, Xiaoxi Mao
Abstract
In this work, we demonstrate that the contextualized word vectors derived from pretrained masked language model-based encoders share a common, perhaps undesirable pattern across layers. Namely, we find cases of persistent outlier neurons within BERT and RoBERTa's hidden state vectors that consistently bear the smallest or largest values in said vectors. In an attempt to investigate the source of this information, we introduce a neuron-level analysis method, which reveals that the outliers are closely related to information captured by positional embeddings. We also pre-train the RoBERTa-base models from scratch and find that the outliers disappear without using positional embeddings. These outliers, we find, are the major cause of anisotropy of encoders' raw vector spaces, and clipping them leads to increased similarity across vectors. We demonstrate this in practice by showing that clipped vectors can more accurately distinguish word senses, as well as lead to better sentence embeddings when mean pooling. In three supervised tasks, we find that clipping does not affect the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fd19fa9-f33f-4af9-832d-fe4e9d4e7ad7Cited by top-tier papers13
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 315 citations
- Outlier Suppression: Pushing the Limit of Low-bit Transformer Language ModelsXiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong et al.NeurIPS 2022 · 238 citations
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 59 citations
Builds on4
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 138 citations
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 5 citations
Related papers
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 29 citations
- On the Robustness of Language Encoders against Grammatical ErrorsFan Yin, Quanyu Long, Tao Meng, Kai-Wei ChangACL 2020 · 32 citations
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 358 citations
- Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language RepresentationsRobert Wolfe, Aylin CaliskanACL 2022 · 16 citations
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang et al.ICLR 2021 · 129 citations
