Interpretable Text Embeddings and Text Similarity Explanation: A Survey
Juri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó, Simon Clematide
摘要
Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search. However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them. In this work, we provide a structured overview of methods specializing in inherently interpretable text embeddings and text similarity explanation, an underexplored research area. We characterize the main ideas, approaches, and trade-offs. We compare means of evaluation, discuss overarching lessons learned and finally identify opportunities and open challenges for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Sentence Smith: Controllable Edits for Evaluating Text EmbeddingsHongji Li, Andrianos Michail, Reto Gubelmann, Simon Clematide 等EMNLP 2025 · 被引用 1 次
- Mapping the Circumplex of Affect: Geometric Analysis of Emotion Representations via Hyperspherical Contrastive LearningYusuke Yamauchi, Akiko AizawaACL 2026
它引用的顶会 Paper16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 被引用 199 次
相关 Paper
- Interpreting BERT-based Text Similarity via Activation and Saliency MapsItzik Malkiel, Dvir Ginzburg, Oren Barkan, Avi Caciularu 等WWW 2022 · 被引用 28 次
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 被引用 5 次
- Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesBin Wang, C.-C. Jay Kuo, Haizhou LiACL 2022 · 被引用 33 次
- Idiosyncrasies in Large Language ModelsMingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter 等ICML 2025
- When is an Embedding Model More Promising than Another?Maxime Darrin, Philippe Formont, Ismail Ben Ayed, Jackie CK Cheung 等NeurIPS 2024 · 被引用 10 次
