MorphTE: Injecting Morphology in Tensorized Embeddings
Guobing Gan, Peng Zhang, Sunzhu Li, Xiuqing Lu, Benyou Wang
摘要
In the era of deep learning, word embeddings are essential when dealing with text tasks. However, storing and accessing these embeddings requires a large amount of space. This is not conducive to the deployment of these models on resource-limited devices. Combining the powerful compression capability of tensor products, we propose a word embedding compression method with morphological augmentation, Morphologically-enhanced Tensorized Embeddings (MorphTE). A word consists of one or more morphemes, the smallest units that bear meaning or have a grammatical function. MorphTE represents a word embedding as an entangled form of its morpheme vectors via the tensor product, which injects prior semantic and grammatical knowledge into the learning of embeddings. Furthermore, the dimensionality of the morpheme vector and the number of morphemes are much smaller than those of words, which greatly reduces the parameters of the word embeddings. We conduct experiments on tasks such as machine translation and question answering. Experimental results on four translation datasets of different languages show that MorphTE can compress word embedding parameters by about 20 times without performance loss and significantly outperforms related embedding compression methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- word2ket: Space-efficient Word Embeddings inspired by Quantum EntanglementAliakbar Panahi, Seyran Saeedi, Tomasz ArodzICLR 2020 · 被引用 44 次
相关 Paper
- From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesKrtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau 等AAAI 2022 · 被引用 3 次
- Accelerating Neural Machine Translation with Partial Word Embedding CompressionFan Zhang, Mei Tu, Jinyao YanAAAI 2021 · 被引用 3 次
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting 等EMNLP 2024 · 被引用 2 次
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv 等EMNLP 2022 · 被引用 3 次
- A Latent Morphology Model for Open-Vocabulary Neural Machine TranslationDuygu Ataman, Wilker Aziz, Alexandra BirchICLR 2020 · 被引用 18 次
