HyperMLP: An Integrated Perspective for Sequence Modeling
Jiecheng Lu, Shihao Yang
Abstract
Self-attention is often viewed as probabilistic query-key lookup, motivating designs that preserve normalized attention scores and fixed positional semantics. We advocate a simpler and more unified perspective: an autoregressive attention head can be viewed as a dynamic two-layer MLP whose weights are instantiated from the context history. From this view, attention scores form an ever-growing hidden representation, and standard MLP activations such as ReLU or GLU naturally implement input-conditioned selection over a context-dependent memory pool rather than a probability distribution. Based on this formulation, we introduce HyperMLP and HyperGLU, which learn dynamic mixing in both feature space and sequence space, using a reverse-offset (lag) layout to align temporal mixing with autoregressive semantics. We provide theoretical characterizations of the expressivity and implications of this structure, and empirically show that Hyper-MLP/HyperGLU consistently outperform strong softmax-attention baselines under matched parameter budgets. Code is available at this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a670a515-3809-4ac0-b8fc-e87b6c0ef53fBuilds on20
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
Related papers
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
- HyperMixer: An MLP-based Low Cost Alternative to TransformersFlorian Mai, Arnaud Pannatier, Fabio Fehr, Haolin Chen et al.ACL 2023 · 13 citations
- JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and AttentionYuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen et al.ICLR 2024 · 49 citations
- Native Hybrid Attention for Efficient Sequence ModelingJusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun et al.ACL 2026 · 8 citations
- Universal Approximation with Softmax AttentionJerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu et al.ICML 2026 · 9 citations
