Isotropy in the Contextual Embedding Space: Clusters and Manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth Church
Abstract
The geometric properties of contextual embedding spaces for deep language models such as BERT and ERNIE, have attracted considerable attention in recent years. Investigations on the contextual embeddings demonstrate a strong anisotropic space such that most of the vectors fall within a narrow cone, leading to high cosine similarities. It is surprising that these LMs are as successful as they are, given that most of their embedding vectors are as similar to one another as they are. In this paper, we argue that the isotropy indeed exists in the space, from a different but more constructive perspective. We identify isolated clusters and low dimensional manifolds in the contextual embedding space, and introduce tools to both qualitatively and quantitatively analyze them. We hope the study in this paper could provide insights towards a better understanding of the deep language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 131dcc6b-e81d-49cb-bad8-3492a3b2d52aCited by top-tier papers47
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 85 citations
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 59 citations
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2024 · 54 citations
- InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingJunda Wu, Tong Yu, Rui Wang, Zhao Song et al.NeurIPS 2023 · 48 citations
- Sharpness-Aware Minimization Leads to Low-Rank FeaturesMaksym Andriushchenko, Dara Bahri, Hossein Mobahi, Nicolas FlammarionNeurIPS 2023 · 48 citations
Related papers
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie et al.ICLR 2021 · 19 citations
- Stable Anisotropic RegularizationWilliam Rudman, Carsten EickhoffICLR 2024 · 13 citations
- Symmetries in language statistics shape the geometry of model representationsDhruva Karkada, Daniel Korchinski, Andres Nava, Matthieu Wyart et al.ICML 2026 · 15 citations
- Token Embeddings Violate the Manifold HypothesisMichael Robinson, Sourya Dey, Tony ChiangNeurIPS 2025 · 19 citations
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson et al.ICML 2020 · 52 citations
