Tuformer: Data-driven Design of Transformers for Improved Generalization or Efficiency
Xiaoyu Liu, Jiahao Su, Furong Huang
摘要
Transformers are neural network architectures that achieve remarkable performance in many areas. However, the core component of Transformers, multi-head self-attention (MHSA), is mainly derived from heuristics, and the interactions across its components are not well understood. To address the problem, we first introduce a mathematically rigorous and yet intuitive tensor diagram representation of MHSA. Guided by tensor diagram representations, we propose a novel design, namely Tunable Transformers (Tuformers), by allowing data-driven weights across heads, whereas MHSA adopts pre-defined and fixed weights across heads, as will be explained in our paper. Tuformers naturally reveal a flexible design space that a user, depending on the needs, can choose a structure that has either improved performance (generalization error) or higher model efficiency. Any pre-trained Transformer can be an initialization of the corresponding Tuformer with trainable number of heads for efficient training and fine-tuning. Tuformers universally outperform Transformers on various tasks across multiple domains under a wide range of model sizes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 被引用 115 次
- Devil is in the Uniformity: Exploring Diverse Learners Within Transformer for Image RestorationShihao Zhou, Dayu Li, Jinshan Pan, Juncheng Zhou 等ICCV 2025 · 被引用 8 次
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 被引用 8 次
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 被引用 7 次
相关 Paper
- Attention-Only Transformers via Unrolled Subspace DenoisingPeng Wang, Yifu Lu, Yaodong Yu, Druv Pai 等ICML 2025
- TransforLearn: Interactive Visual Tutorial for the Transformer ModelLin Gao, Zekai Shao, Ziqin Luo, Haibo Hu 等IEEE VIS 2023 · 被引用 13 次
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- Improving Transformer with an Admixture of Attention HeadsTan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen 等NeurIPS 2022 · 被引用 38 次
- Compositional Attention: Disentangling Search and RetrievalSarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio 等ICLR 2022 · 被引用 20 次
