Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
Yekun Chai, Jin Shuo, Xinwen Hou
Abstract
Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multiheaded dot product attention by attending to all the global contexts at different locations. Through a pseudo information highway, we introduce a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations. The subsidiary content-based SDU gates allow for the information flow of modulated latent embeddings through skipped connections, leading to a clear margin of convergence speed with gradient descent algorithms. We may unveil the role of gating mechanism to aid in the contextbased Transformer modules, with hypothesizing that SDU gates, especially on shallow layers, could push it faster to step towards suboptimal points during the optimization process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a6542e0-2bf2-47b6-b990-b1db6326878aCited by top-tier papers5
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao et al.ICML 2026 · 72 citations
- KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual MatricesWuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo MandicICML 2026 · 8 citations
- Autoregressive Pre-Training on Pixels and TextsYekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang et al.EMNLP 2024 · 1 citation
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng et al.ICLR 2025
- Residual Matrix Transformers: Scaling the Size of the Residual StreamBrian Mak, Jeffrey FlaniganICML 2025
Builds on1
Related papers
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 47 citations
- Sparse Modular Activation for Efficient Sequence ModelingLiliang Ren, Yang Liu, Shuohang Wang, Yichong Xu et al.NeurIPS 2023 · 23 citations
- Long Range Language Modeling via Gated State SpacesHarsh Mehta, Ankit Gupta, Ashok Cutkosky, Behnam NeyshaburICLR 2023 · 45 citations
- Dual-Granularity Memory for Efficient Video GenerationHongjun Wang, Lin Liu, Jianguo Li, Tao LinCVPR 2026
