HyperMixer: An MLP-based Low Cost Alternative to Transformers
Florian Mai, Arnaud Pannatier, Fabio Fehr, Haolin Chen, François Marelli, François Fleuret, James Henderson
摘要
Transformer-based architectures are the model of choice for natural language understanding, but they come at a significant cost, as they have quadratic complexity in the input length, require a lot of training data, and can be difficult to tune. In the pursuit of lower costs, we investigate simple MLP-based architectures. We find that existing architectures such as MLPMixer, which achieves token mixing through a static MLP applied to each feature independently, are too detached from the inductive biases required for natural language understanding. In this paper, we propose a simple variant, HyperMixer, which forms the token mixing MLP dynamically using hypernetworks. Empirically, we demonstrate that our model performs better than alternative MLP-based models, and on par with Transformers. In contrast to Transformers, HyperMixer achieves these results at substantially lower costs in terms of processing time, training data, and hyperparameter tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 被引用 912 次
- MAXIM: Multi-Axis MLP for Image ProcessingZhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang 等CVPR 2022 · 被引用 550 次
- AS-MLP: An Axial Shifted MLP Architecture for VisionDongze Lian, Zehao Yu, Xing Sun, Shenghua GaoICLR 2022 · 被引用 217 次
相关 Paper
- DynaMixer: A Vision MLP Architecture with Dynamic MixingZiyu Wang, Wenhao Jiang, Yiming Zhu, Li Yuan 等ICML 2022 · 被引用 55 次
- StockMixer: A Simple Yet Strong MLP-Based Architecture for Stock Price ForecastingJinyong Fan, Yanyan ShenAAAI 2024 · 被引用 43 次
- FFT-Based Dynamic Token Mixer for VisionYuki Tatsunami, Masato TakiAAAI 2024 · 被引用 73 次
- The Unstoppable Rise of Computational Linguistics in Deep LearningJames HendersonACL 2020 · 被引用 4 次
- MLPs Learn In-Context on Regression and Classification TasksWilliam Lingxiao Tong, Cengiz PehlevanICLR 2025
