HashFormers: Towards Vocabulary-independent Pre-trained Transformers
Huiyin Xue, Nikolaos Aletras
Abstract
Transformer-based pre-trained language models are vocabulary-dependent, mapping by default each token to its corresponding embedding. This one-to-one mapping results into embedding matrices that occupy a lot of memory (i.e. millions of parameters) and grow linearly with the size of the vocabulary. Previous work on on-device transformers dynamically generate token embeddings on-the-fly without embedding matrices using locality-sensitive hashing over morphological information. These embeddings are subsequently fed into transformer layers for text classification. However, these methods are not pre-trained. Inspired by this line of work, we propose HASHFORMERS, a new family of vocabulary-independent pretrained transformers that support an unlimited vocabulary (i.e. all possible tokens in a corpus) given a substantially smaller fixed-sized embedding matrix. We achieve this by first introducing computationally cheap hashing functions that bucket together individual tokens to embeddings. We also propose three variants that do not require an embedding matrix at all, further reducing the memory requirements. We empirically demonstrate that HASHFORM-ERS are more memory efficient compared to standard pre-trained transformers while achieving comparable predictive performance when fine-tuned on multiple text classification tasks. For example, our most efficient HASHFORMER variant has a negligible performance degradation (0.4% on GLUE) using only 99.1K parameters for representing the embeddings compared to 12.3-38M parameters of state-of-the-art models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 422c2a73-ef00-4d77-bc55-89b1e4ede4edCited by top-tier papers2
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting et al.EMNLP 2024 · 2 citations
- Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?Ahmed Alajrami, Katerina Margatina, Nikolaos AletrasEMNLP 2023 · 1 citation
Builds on3
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson et al.ICLR 2021 · 11 citations
Related papers
- MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected LayersNing Ding, Yehui Tang, Haochen Qin, Zhenli Zhou et al.NeurIPS 2024 · 8 citations
- Improving Transformer with an Admixture of Attention HeadsTan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen et al.NeurIPS 2022 · 38 citations
- LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language ModelsHaoyu Wang, Ruirui Li, Haoming Jiang, Zhengyang Wang et al.KDD 2023 · 5 citations
- Sparse Attention with Learning to HashZhiqing Sun, Yiming Yang, Shinjae YooICLR 2022 · 21 citations
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
