TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
Haiyang Wang, Yue Fan, Muhammad Ferjad Naeem, Yongqin Xian, Jan Eric Lenssen, Liwei Wang, Federico Tombari, Bernt Schiele
摘要
Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises primarily from their dependence on a fixed number of parameters within linear projections. When architectural modifications (e.g., channel dimensions) are introduced, the entire model typically requires retraining from scratch. As model sizes continue growing, this strategy results in increasingly high computational costs and becomes unsustainable. To overcome this problem, we introduce Tokenformer, a natively scalable architecture that leverages the attention mechanism not only for computations among input tokens but also for interactions between tokens and model parameters, thereby enhancing architectural flexibility. By treating model parameters as tokens, we replace all the linear projections in Transformers with our token-parameter attention layer, where input tokens act as queries and model parameters as keys and values. This reformulation allows for progressive and efficient scaling without necessitating retraining from scratch. Our model scales from 124M to 1.4B parameters by incrementally adding new key-value parameter pairs, achieving performance comparable to Transformers trained from scratch while greatly reducing training costs. Code and models are available at https://github.com/Haiyang-W/TokenFormer.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language InterfaceHao Tang, Chen-Wei Xie, Haiyang Wang, Xiaoyi Bao 等NeurIPS 2025 · 被引用 30 次
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
相关 Paper
- Progressive Token Length Scaling in Transformer Encoders for Efficient Universal SegmentationAbhishek Aich, Yumin Suh, Samuel Schulter, Manmohan ChandrakerICLR 2025
- Fcaformer: Forward Cross Attention in Hybrid Vision TransformerHaokui Zhang, Wenze Hu, Xiaoyu WangICCV 2023 · 被引用 10 次
- Inner-layer Token Self-modulation as Another Scaling Axis for LLMsYebin Yang, Huaijin Wu, Jingtao Han, Yu Wang 等ICML 2026 · 被引用 1 次
- Low-Rank Bottleneck in Multi-head Attention ModelsSrinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi 等ICML 2020 · 被引用 130 次
- Over-Tokenized Transformer: Vocabulary is Generally Worth ScalingHongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng 等ICML 2025
