Interpretable Text Embeddings and Text Similarity Explanation: A Survey
Juri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó, Simon Clematide
Abstract
Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search. However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them. In this work, we provide a structured overview of methods specializing in inherently interpretable text embeddings and text similarity explanation, an underexplored research area. We characterize the main ideas, approaches, and trade-offs. We compare means of evaluation, discuss overarching lessons learned and finally identify opportunities and open challenges for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94b1be13-0d56-4429-a218-cd1298c35992Cited by top-tier papers2
- Sentence Smith: Controllable Edits for Evaluating Text EmbeddingsHongji Li, Andrianos Michail, Reto Gubelmann, Simon Clematide et al.EMNLP 2025 · 1 citation
- Mapping the Circumplex of Affect: Geometric Analysis of Emotion Representations via Hyperspherical Contrastive LearningYusuke Yamauchi, Akiko AizawaACL 2026
Builds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
Related papers
- Interpreting BERT-based Text Similarity via Activation and Saliency MapsItzik Malkiel, Dvir Ginzburg, Oren Barkan, Avi Caciularu et al.WWW 2022 · 28 citations
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 5 citations
- Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesBin Wang, C.-C. Jay Kuo, Haizhou LiACL 2022 · 33 citations
- Idiosyncrasies in Large Language ModelsMingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter et al.ICML 2025
- When is an Embedding Model More Promising than Another?Maxime Darrin, Philippe Formont, Ismail Ben Ayed, Jackie CK Cheung et al.NeurIPS 2024 · 10 citations
