Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
Yekun Chai, Jin Shuo, Xinwen Hou
摘要
Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multiheaded dot product attention by attending to all the global contexts at different locations. Through a pseudo information highway, we introduce a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations. The subsidiary content-based SDU gates allow for the information flow of modulated latent embeddings through skipped connections, leading to a clear margin of convergence speed with gradient descent algorithms. We may unveil the role of gating mechanism to aid in the contextbased Transformer modules, with hypothesizing that SDU gates, especially on shallow layers, could push it faster to step towards suboptimal points during the optimization process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- mHC: Manifold-Constrained Hyper-ConnectionsZhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao 等ICML 2026 · 被引用 72 次
- KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual MatricesWuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo MandicICML 2026 · 被引用 8 次
- Autoregressive Pre-Training on Pixels and TextsYekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang 等EMNLP 2024 · 被引用 1 次
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng 等ICLR 2025
- Residual Matrix Transformers: Scaling the Size of the Residual StreamBrian Mak, Jeffrey FlaniganICML 2025
它引用的顶会 Paper1
相关 Paper
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 被引用 47 次
- Sparse Modular Activation for Efficient Sequence ModelingLiliang Ren, Yang Liu, Shuohang Wang, Yichong Xu 等NeurIPS 2023 · 被引用 23 次
- Long Range Language Modeling via Gated State SpacesHarsh Mehta, Ankit Gupta, Ashok Cutkosky, Behnam NeyshaburICLR 2023 · 被引用 45 次
- Dual-Granularity Memory for Efficient Video GenerationHongjun Wang, Lin Liu, Jianguo Li, Tao LinCVPR 2026
