Interpreting Embedding Spaces by Conceptualization
Adi Simhi, Shaul Markovitch
Abstract
One of the main methods for computational interpretation of a text is mapping it into a vector in some embedding space. Such vectors can then be used for a variety of textual processing tasks. Recently, most embedding spaces are a product of training large language models (LLMs). One major drawback of this type of representation is their incomprehensibility to humans. Understanding the embedding space is crucial for several important needs, including the need to debug the embedding method and compare it to alternatives, and the need to detect biases hidden in the model. In this paper, we present a novel method of understanding embeddings by transforming a latent embedding space into a comprehensible conceptual space. We present an algorithm for deriving a conceptual space with dynamic on-demand granularity. We devise a new evaluation method, using either human rater or LLM-based raters, to show that the conceptualized vectors indeed represent the semantics of the original latent ones. We show the use of our method for various tasks, including comparing the semantics of alternative models and tracing the layers of the LLM. The code is available online https://github.com/adiSimhi/Interpreting-Embedding-Spaces-by-Conceptualization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8c26621-d493-4817-81ed-b929b02ec0dfCited by top-tier papers5
- SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringXiaopeng Li, Shasha Li, Shezheng Song, Huijun Liu et al.AAAI 2025 · 11 citations
- Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceShahar Katz, Yonatan Belinkov, Mor Geva, Lior WolfEMNLP 2024 · 2 citations
- PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderYiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, Jun YuACL 2025 · 2 citations
- EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsMohammadamin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong et al.ICML 2026 · 1 citation
- A General Framework for Producing Interpretable Semantic Text EmbeddingsYiqun Sun, Qiang Huang, Yixuan Tang, Anthony Kum Hoe Tung et al.ICLR 2025
Builds on6
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word EmbeddingsBinny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus StrohmaierWWW 2020 · 40 citations
- Analyzing Transformers in Embedding SpaceGuy Dar, Mor Geva, Ankit Gupta, Jonathan BerantACL 2023 · 36 citations
- Sparsity Makes Sense: Word Sense Disambiguation Using Sparse Contextualized Word RepresentationsGábor BerendEMNLP 2020 · 18 citations
Related papers
- Demystifying Embedding Spaces using Large Language ModelsGuy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Jihwan Jeong et al.ICLR 2024 · 24 citations
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha et al.ICML 2026 · 6 citations
- Interpretability of Language Models via Task SpacesLucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke HupkesACL 2024
