Stable Anisotropic Regularization
William Rudman, Carsten Eickhoff
摘要
Given the success of Large Language Models (LLMs), there has been considerable interest in studying the properties of model activations. The literature overwhelmingly agrees that LLM representations are dominated by a few"outlier dimensions"with exceedingly high variance and magnitude. Several studies in Natural Language Processing (NLP) have sought to mitigate the impact of such outlier dimensions and force LLMs to be isotropic (i.e., have uniform variance across all dimensions in embedding space). Isotropy is thought to be a desirable property for LLMs that improves model performance and more closely aligns textual representations with human intuition. However, many of the claims regarding isotropy in NLP have been based on the average cosine similarity of embeddings, which has recently been shown to be a flawed measure of isotropy. In this paper, we propose I-STAR: IsoScore*-based STable Anisotropic Regularization, a novel regularization method that can be used to increase or decrease levels of isotropy in embedding space during training. I-STAR uses IsoScore*, the first accurate measure of isotropy that is both differentiable and stable on mini-batch computations. In contrast to several previous works, we find that decreasing isotropy in contextualized embeddings improves performance on the majority of tasks and models considered in this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal InputsMustafa Shukor, Matthieu CordNeurIPS 2024 · 被引用 27 次
- Metis: Training LLMs with FP4 QuantizationHengjie Cao, Mengyi Chen, Yifeng Yang, Fang Dong 等ICLR 2026 · 被引用 10 次
- Rare Text Semantics Were Always There in Your Diffusion TransformerSeil Kang, Woojung Han, Dayun Ju, Seong Jae HwangNeurIPS 2025 · 被引用 6 次
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMsZhongyang Li, Ziyue Li, Tianyi ZhouICLR 2026 · 被引用 5 次
- Better Embeddings with Coupled AdamFelix Stollenwerk, Tobias StollenwerkACL 2025 · 被引用 4 次
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 被引用 59 次
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 被引用 50 次
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 被引用 29 次
- Embedding Compression with Isotropic Iterative QuantizationSiyu Liao, Jie Chen, Yanzhi Wang, Qinru Qiu 等AAAI 2020 · 被引用 16 次
相关 Paper
- Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text GenerationDhrupad Bhardwaj, Julia Kempe, Tim G. J. RudnerICML 2026
- Revisiting Anisotropy in Language Transformers: The Geometry of Learning DynamicsRaphael Bernas, Fanny Jourdan, Antonin Poché, Céline HudelotICML 2026 · 被引用 3 次
- Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee 等ACL 2022 · 被引用 42 次
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan 等ICML 2026 · 被引用 5 次
- AoE: Angle-optimized Embeddings for Semantic Textual SimilarityXianming Li, Jing LiACL 2024 · 被引用 22 次
