Stable Anisotropic Regularization
William Rudman, Carsten Eickhoff
Abstract
Given the success of Large Language Models (LLMs), there has been considerable interest in studying the properties of model activations. The literature overwhelmingly agrees that LLM representations are dominated by a few"outlier dimensions"with exceedingly high variance and magnitude. Several studies in Natural Language Processing (NLP) have sought to mitigate the impact of such outlier dimensions and force LLMs to be isotropic (i.e., have uniform variance across all dimensions in embedding space). Isotropy is thought to be a desirable property for LLMs that improves model performance and more closely aligns textual representations with human intuition. However, many of the claims regarding isotropy in NLP have been based on the average cosine similarity of embeddings, which has recently been shown to be a flawed measure of isotropy. In this paper, we propose I-STAR: IsoScore*-based STable Anisotropic Regularization, a novel regularization method that can be used to increase or decrease levels of isotropy in embedding space during training. I-STAR uses IsoScore*, the first accurate measure of isotropy that is both differentiable and stable on mini-batch computations. In contrast to several previous works, we find that decreasing isotropy in contextualized embeddings improves performance on the majority of tasks and models considered in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc6f46a7-1d01-4016-a1ad-71f7b5ccb210Cited by top-tier papers10
- Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal InputsMustafa Shukor, Matthieu CordNeurIPS 2024 · 27 citations
- Metis: Training LLMs with FP4 QuantizationHengjie Cao, Mengyi Chen, Yifeng Yang, Fang Dong et al.ICLR 2026 · 10 citations
- Rare Text Semantics Were Always There in Your Diffusion TransformerSeil Kang, Woojung Han, Dayun Ju, Seong Jae HwangNeurIPS 2025 · 6 citations
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMsZhongyang Li, Ziyue Li, Tianyi ZhouICLR 2026 · 5 citations
- Better Embeddings with Coupled AdamFelix Stollenwerk, Tobias StollenwerkACL 2025 · 4 citations
Builds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 59 citations
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 50 citations
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 29 citations
- Embedding Compression with Isotropic Iterative QuantizationSiyu Liao, Jie Chen, Yanzhi Wang, Qinru Qiu et al.AAAI 2020 · 16 citations
Related papers
- Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text GenerationDhrupad Bhardwaj, Julia Kempe, Tim G. J. RudnerICML 2026
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 3 citations
- Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee et al.ACL 2022 · 42 citations
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan et al.ICML 2026 · 5 citations
- AoE: Angle-optimized Embeddings for Semantic Textual SimilarityXianming Li, Jing LiACL 2024 · 22 citations
