Demystifying Embedding Spaces using Large Language Models
Guy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Jihwan Jeong, Lior Shani, Azamat Tulepbergenov, Deepak Ramachandran, Martin Mladenov, Craig Boutilier
Abstract
Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use of these compressed representations, meaningful interpretation usually requires visualization using dimensionality reduction or specialized machine learning interpretability methods. This paper addresses the challenge of making such embeddings more interpretable and broadly useful, by employing Large Language Models (LLMs) to directly interact with embeddings -- transforming abstract vectors into understandable narratives. By injecting embeddings into LLMs, we enable querying and exploration of complex embedding data. We demonstrate our approach on a variety of diverse tasks, including: enhancing concept activation vectors (CAVs), communicating novel embedded entities, and decoding user preferences in recommender systems. Our work couples the immense information potential of embeddings with the interpretative power of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dde6b9d-7947-4ff9-a431-6d3db22eaf05Cited by top-tier papers8
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual ExplanationsLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang et al.ICML 2024 · 18 citations
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin et al.NeurIPS 2025 · 7 citations
- Embedding-Aligned Language ModelsGuy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Lior Shani et al.NeurIPS 2024 · 7 citations
- EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsMohammadamin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong et al.ICML 2026 · 1 citation
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation ModelWei-Hsin Yeh, Yu-An Su, Chih-Ning Chen, Yi-Hsueh Lin et al.ACL 2025 · 1 citation
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
Related papers
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
- Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or WorseAaron Imani, Mohammad Moshirpour, Iftekhar AhmedICSE 2026
- Interpreting Embedding Spaces by ConceptualizationAdi Simhi, Shaul MarkovitchEMNLP 2023 · 6 citations
- LLMRG: Improving Recommendations through Large Language Model Reasoning GraphsYan Wang, Zhixuan Chu, Xin Ouyang, Simeng Wang et al.AAAI 2024 · 47 citations
- Bridging LLM Embeddings and VAE Parameters for Disentangled RecommendationNhu-Thuat Tran, Hady W. LauwSIGIR 2026
