Uncovering Meanings of Embeddings via Partial Orthogonality
Yibo Jiang, Bryon Aragam, Victor Veitch
Abstract
Machine learning tools often rely on embedding text as vectors of real numbers. In this paper, we study how the semantic structure of language is encoded in the algebraic structure of such embeddings. Specifically, we look at a notion of semantic independence'' capturing the idea that, e.g., eggplant'' and tomato'' are independent given vegetable''. Although such examples are intuitive, it is difficult to formalize such a notion of semantic independence. The key observation here is that any sensible formalization should obey a set of so-called independence axioms, and thus any algebraic encoding of this structure should also obey these axioms. This leads us naturally to use partial orthogonality as the relevant algebraic structure. We develop theory and methods that allow us to demonstrate that partial orthogonality does indeed capture semantic independence. Complementary to this, we also introduce the concept of independence preserving embeddings where embeddings preserve the conditional independence structures of a distribution, and we prove the existence of such embeddings and approximations to them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam et al.ICML 2024 · 68 citations
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf et al.NeurIPS 2024 · 37 citations
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou et al.ICLR 2026 · 29 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Robust Learning with the Hilbert-Schmidt Independence CriterionDaniel Greenfeld, Uri ShalitICML 2020 · 73 citations
- Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsMatthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille et al.ICCV 2023 · 51 citations
- Efficient Bayesian network structure learning via local Markov boundary searchMing Gao, Bryon AragamNeurIPS 2021 · 20 citations
Related papers
- Discovering Universal Geometry in Embeddings with ICAHiroaki Yamagiwa, Momose Oyama, Hidetoshi ShimodairaEMNLP 2023 · 6 citations
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 10 citations
- Transformers learn factored representationsAdam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow et al.ICML 2026 · 2 citations
- Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language ModelsNa Li, Hanane Kteich, Zied Bouraoui, Steven SchockaertSIGIR 2023 · 3 citations
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon et al.ACL 2020 · 66 citations
