Obtaining Better Static Word Embeddings Using Contextual Embedding Models
Prakhar Gupta, Martin Jaggi
摘要
The advent of contextual word embeddingsrepresentations of words which incorporate semantic and syntactic information from their context-has led to tremendous improvements on a wide variety of NLP tasks. However, recent contextual models have prohibitively high computational cost in many use-cases and are often hard to interpret. In this work, we demonstrate that our proposed distillation method, which is a simple extension of CBOW-based training, allows to significantly improve computational efficiency of NLP applications, while outperforming the quality of existing static embeddings trained from scratch as well as those distilled from previously proposed methods. As a side-effect, our approach also allows a fair comparison of both contextual and static embeddings via standard lexical evaluation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Backpack Language ModelsJohn Hewitt, John Thickstun, Christopher D. Manning, Percy LiangACL 2023 · 被引用 9 次
- Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language ModelsNa Li, Hanane Kteich, Zied Bouraoui, Steven SchockaertSIGIR 2023 · 被引用 3 次
- DiVa: An Iterative Framework to Harvest More Diverse and Valid Labels from User Comments for MusicHongru Liang, Jingyao Liu, Yuanxin Xiang, Jiachen Du 等ACM MM 2023
- Imputing Out-of-Vocabulary Embeddings with LOVE Makes LanguageModels Robust with Little CostLihu Chen, Gaël Varoquaux, Fabian M. SuchanekACL 2022
它引用的顶会 Paper3
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 被引用 137 次
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 被引用 7 次
相关 Paper
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima 等EMNLP 2025 · 被引用 1 次
- Using Context-to-Vector with Graph Retrofitting to Improve Word EmbeddingsJiangbin Zheng, Yile Wang, Ge Wang, Jun Xia 等ACL 2022 · 被引用 27 次
- Scalable Attentive Sentence Pair Modeling via Distilled Sentence EmbeddingOren Barkan, Noam Razin, Itzik Malkiel, Ori Katz 等AAAI 2020 · 被引用 37 次
- Token Distillation: Attention-Aware Input Embeddings for New TokensKonstantin Dobler, Desmond Elliott, Gerard de MeloICLR 2026 · 被引用 7 次
- Distilling Linguistic Context for Language Model CompressionGeondo Park, Gyeongman Kim, Eunho YangEMNLP 2021 · 被引用 24 次
