Obtaining Better Static Word Embeddings Using Contextual Embedding Models
Prakhar Gupta, Martin Jaggi
Abstract
The advent of contextual word embeddingsrepresentations of words which incorporate semantic and syntactic information from their context-has led to tremendous improvements on a wide variety of NLP tasks. However, recent contextual models have prohibitively high computational cost in many use-cases and are often hard to interpret. In this work, we demonstrate that our proposed distillation method, which is a simple extension of CBOW-based training, allows to significantly improve computational efficiency of NLP applications, while outperforming the quality of existing static embeddings trained from scratch as well as those distilled from previously proposed methods. As a side-effect, our approach also allows a fair comparison of both contextual and static embeddings via standard lexical evaluation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7caecbb1-3c1e-4940-8dd3-2780bc969c28Cited by top-tier papers4
- Backpack Language ModelsJohn Hewitt, John Thickstun, Christopher D. Manning, Percy LiangACL 2023 · 9 citations
- Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language ModelsNa Li, Hanane Kteich, Zied Bouraoui, Steven SchockaertSIGIR 2023 · 3 citations
- DiVa: An Iterative Framework to Harvest More Diverse and Valid Labels from User Comments for MusicHongru Liang, Jingyao Liu, Yuanxin Xiang, Jiachen Du et al.ACM MM 2023
- Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostLihu Chen, Gaël Varoquaux, Fabian M. SuchanekACL 2022
Builds on3
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 137 citations
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 7 citations
Related papers
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima et al.EMNLP 2025 · 1 citation
- Using Context-to-Vector with Graph Retrofitting to Improve Word EmbeddingsJiangbin Zheng, Yile Wang, Ge Wang, Jun Xia et al.ACL 2022 · 27 citations
- Scalable Attentive Sentence Pair Modeling via Distilled Sentence EmbeddingOren Barkan, Noam Razin, Itzik Malkiel, Ori Katz et al.AAAI 2020 · 37 citations
- Token Distillation: Attention-Aware Input Embeddings for New TokensKonstantin Dobler, Desmond Elliott, Gerard de MeloICLR 2026 · 7 citations
- Distilling Linguistic Context for Language Model CompressionGeondo Park, Gyeongman Kim, Eunho YangEMNLP 2021 · 24 citations
