Interpreting Embedding Spaces by Conceptualization
Adi Simhi, Shaul Markovitch
摘要
One of the main methods for computational interpretation of a text is mapping it into a vector in some embedding space. Such vectors can then be used for a variety of textual processing tasks. Recently, most embedding spaces are a product of training large language models (LLMs). One major drawback of this type of representation is their incomprehensibility to humans. Understanding the embedding space is crucial for several important needs, including the need to debug the embedding method and compare it to alternatives, and the need to detect biases hidden in the model. In this paper, we present a novel method of understanding embeddings by transforming a latent embedding space into a comprehensible conceptual space. We present an algorithm for deriving a conceptual space with dynamic on-demand granularity. We devise a new evaluation method, using either human rater or LLM-based raters, to show that the conceptualized vectors indeed represent the semantics of the original latent ones. We show the use of our method for various tasks, including comparing the semantics of alternative models and tracing the layers of the LLM. The code is available online https://github.com/adiSimhi/Interpreting-Embedding-Spaces-by-Conceptualization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringXiaopeng Li, Shasha Li, Shezheng Song, Huijun Liu 等AAAI 2025 · 被引用 11 次
- Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceShahar Katz, Yonatan Belinkov, Mor Geva, Lior WolfEMNLP 2024 · 被引用 2 次
- PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderYiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, Jun YuACL 2025 · 被引用 2 次
- EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsMohammadamin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong 等ICML 2026 · 被引用 1 次
- A General Framework for Producing Interpretable Semantic Text EmbeddingsYiqun Sun, Qiang Huang, Yixuan Tang, Anthony Kum Hoe Tung 等ICLR 2025
它引用的顶会 Paper6
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word EmbeddingsBinny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus StrohmaierWWW 2020 · 被引用 40 次
- Analyzing Transformers in Embedding SpaceGuy Dar, Mor Geva, Ankit Gupta, Jonathan BerantACL 2023 · 被引用 36 次
- Sparsity Makes Sense: Word Sense Disambiguation Using Sparse Contextualized Word RepresentationsGábor BerendEMNLP 2020 · 被引用 18 次
相关 Paper
- Demystifying Embedding Spaces using Large Language ModelsGuy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Jihwan Jeong 等ICLR 2024 · 被引用 24 次
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello 等NeurIPS 2024 · 被引用 26 次
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha 等ICML 2026 · 被引用 6 次
- Interpretability of Language Models via Task SpacesLucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke HupkesACL 2024
