Which transformer architecture fits my data? A vocabulary bottleneck in self-attention
Noam Wies, Yoav Levine, Daniel Jannai, Amnon Shashua
Abstract
After their successful debut in natural language processing, Transformer architectures are now becoming the de-facto standard in many domains. An obstacle for their deployment over new modalities is the architectural configuration: the optimal depth-to-width ratio has been shown to dramatically vary across data types (e.g., x larger over images than over language). We theoretically predict the existence of an embedding rank bottleneck that limits the contribution of self-attention width to the Transformer expressivity. We thus directly tie the input vocabulary size and rank to the optimal depth-to-width ratio, since a small vocabulary size or rank dictates an added advantage of depth over width. We empirically demonstrate the existence of this bottleneck and its implications on the depth-to-width interplay of Transformer architectures, linking the architecture variability across domains to the often glossed-over usage of different vocabulary sizes or embedding ranks in different domains. As an additional benefit, our rank bottlenecking framework allows us to identify size redundancies of in leading NLP models such as ALBERT and T5.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4ad9c2d-8ad7-4c9e-945d-ff68095cb87aCited by top-tier papers8
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho et al.NeurIPS 2022 · 311 citations
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
- The Inductive Bias of In-Context Learning: Rethinking Pretraining Example DesignYoav Levine, Noam Wies, Daniel Jannai, Dan Navon et al.ICLR 2022 · 43 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
Related papers
- Low-Rank Bottleneck in Multi-head Attention ModelsSrinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi et al.ICML 2020 · 130 citations
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and CompressibilityAnnan Yu, Danielle C. Maddix, Boran Han, Xiyuan Zhang et al.ICLR 2026 · 9 citations
- Training compute-optimal transformer encoder modelsMegi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCunEMNLP 2025 · 1 citation
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsSofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang et al.NeurIPS 2025 · 6 citations
- Go Wider Instead of DeeperFuzhao Xue, Ziji Shi, Futao Wei, Yuxuan Lou et al.AAAI 2022 · 106 citations
