All Word Embeddings from One Embedding
Sho Takase, Sosuke Kobayashi
Abstract
In neural network-based models for natural language processing (NLP), the largest part of the parameters often consists of word embeddings. Conventional models prepare a large embedding matrix whose size depends on the vocabulary size. Therefore, storing these models in memory and disk storage is costly. In this study, to reduce the total number of parameters, the embeddings for all words are represented by transforming a shared embedding. The proposed method, ALONE (all word embeddings from one), constructs the embedding of a word by modifying the shared embedding with a filter vector, which is word-specific but non-trainable. Then, we input the constructed embedding into a feed-forward neural network to increase its expressiveness. Naively, the filter vectors occupy the same memory size as the conventional embedding matrix, which depends on the vocabulary size. To solve this issue, we also introduce a memory-efficient filter construction approach. We indicate our ALONE can be used as word representation sufficiently through an experiment on the reconstruction of pre-trained word embeddings. In addition, we also conduct experiments on NLP application tasks: machine translation and summarization. We combined ALONE with the current state-of-the-art encoderdecoder model, the Transformer [36], and achieved comparable scores on WMT 2014 English-to-German translation and DUC 2004 very short summarization with less parameters 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesKrtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau et al.AAAI 2022 · 3 citations
- DictFormer: Tiny Transformer with Shared DictionaryQian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen et al.ICLR 2022 · 12 citations
- word2ket: Space-efficient Word Embeddings inspired by Quantum EntanglementAliakbar Panahi, Seyran Saeedi, Tomasz ArodzICLR 2020 · 44 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Softmax Output Approximation for Activation Memory-Efficient Training of Attention-based NetworksChanghyeon Lee, Seulki LeeNeurIPS 2023 · 4 citations
