All Word Embeddings from One Embedding
Sho Takase, Sosuke Kobayashi
摘要
In neural network-based models for natural language processing (NLP), the largest part of the parameters often consists of word embeddings. Conventional models prepare a large embedding matrix whose size depends on the vocabulary size. Therefore, storing these models in memory and disk storage is costly. In this study, to reduce the total number of parameters, the embeddings for all words are represented by transforming a shared embedding. The proposed method, ALONE (all word embeddings from one), constructs the embedding of a word by modifying the shared embedding with a filter vector, which is word-specific but non-trainable. Then, we input the constructed embedding into a feed-forward neural network to increase its expressiveness. Naively, the filter vectors occupy the same memory size as the conventional embedding matrix, which depends on the vocabulary size. To solve this issue, we also introduce a memory-efficient filter construction approach. We indicate our ALONE can be used as word representation sufficiently through an experiment on the reconstruction of pre-trained word embeddings. In addition, we also conduct experiments on NLP application tasks: machine translation and summarization. We combined ALONE with the current state-of-the-art encoderdecoder model, the Transformer [36], and achieved comparable scores on WMT 2014 English-to-German translation and DUC 2004 very short summarization with less parameters 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesKrtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau 等AAAI 2022 · 被引用 3 次
- DictFormer: Tiny Transformer with Shared DictionaryQian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen 等ICLR 2022 · 被引用 12 次
- word2ket: Space-efficient Word Embeddings inspired by Quantum EntanglementAliakbar Panahi, Seyran Saeedi, Tomasz ArodzICLR 2020 · 被引用 44 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Softmax Output Approximation for Activation Memory-Efficient Training of Attention-based NetworksChanghyeon Lee, Seulki LeeNeurIPS 2023 · 被引用 4 次
