Discovering Universal Geometry in Embeddings with ICA
Hiroaki Yamagiwa, Momose Oyama, Hidetoshi Shimodaira
Abstract
This study utilizes Independent Component Analysis (ICA) to unveil a consistent semantic structure within embeddings of words or images. Our approach extracts independent semantic components from the embeddings of a pre-trained model by leveraging anisotropic information that remains after the whitening process in Principal Component Analysis (PCA). We demonstrate that each embedding can be expressed as a composition of a few intrinsic interpretable axes and that these semantic axes remain consistent across different languages, algorithms, and modalities. The discovery of a universal semantic structure in the geometric patterns of embeddings enhances our understanding of the representations in embeddings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context LearningDake Bu, Wei Huang, Andi Han, Atsushi Nitanda et al.NeurIPS 2024 · 11 citations
- Plug-and-Play Compositionality for Boosting Continual Learning with Foundation ModelsWeiduo Liao, Fei Han, Hisao Ishibuchi, Qingfu Zhang et al.ICLR 2026
Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 59 citations
- The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word EmbeddingsBinny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus StrohmaierWWW 2020 · 40 citations
Related papers
- Uncovering Meanings of Embeddings via Partial OrthogonalityYibo Jiang, Bryon Aragam, Victor VeitchNeurIPS 2023 · 20 citations
- A Geometric Analysis of Deep Generative Image Models and Its ApplicationsBinxu Wang, Carlos R. PonceICLR 2021 · 42 citations
- LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable DirectionsOguz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er, Pinar YanardagICCV 2021 · 71 citations
- Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsMatthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille et al.ICCV 2023 · 51 citations
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou et al.NeurIPS 2023 · 13 citations
