Extending Multi-Sense Word Embedding to Phrases and Sentences for Unsupervised Semantic Applications
Haw-Shiuan Chang, Amol Agrawal, Andrew McCallum
Abstract
Most unsupervised NLP models represent each word with a single point or single region in semantic space, while the existing multi-sense word embeddings cannot represent longer word sequences like phrases or sentences. We propose a novel embedding method for a text sequence (a phrase or a sentence) where each sequence is represented by a distinct set of multi-mode codebook embeddings to capture different semantic facets of its meaning. The codebook embeddings can be viewed as the cluster centers which summarize the distribution of possibly co-occurring words in a pre-trained word embedding space. We introduce an end-to-end trainable neural model that directly predicts the set of cluster centers from the input text sequence during test time. Our experiments show that the per-sentence codebook embeddings significantly improve the performances in unsupervised sentence similarity and extractive summarization benchmarks. In phrase similarity experiments, we discover that the multi-facet embeddings provide an interpretable semantic representation but do not outperform the single-facet baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 839a8439-863d-4f52-9e57-218f7cd37033Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima et al.EMNLP 2025 · 1 citation
- Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language ModelsNa Li, Hanane Kteich, Zied Bouraoui, Steven SchockaertSIGIR 2023 · 3 citations
- FIRE: Semantic Field of Words Represented as Non-Linear FunctionsXin Du, Kumiko Tanaka-IshiiNeurIPS 2022
- SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain MappingMarc Felix Brinner, Sina ZarrießACL 2026
- Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsJames Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang et al.EMNLP 2023
