EIT: Enhanced Interactive Transformer
Tong Zheng, Bei Li, Huiwen Bao, Tong Xiao, JingBo Zhu
摘要
Two principles: the complementary princi-001 ple and the consensus principle are widely 002 acknowledged in the literature of multi-view 003 learning. However, the current design of Multi-004 head self-attention, an instance of multi-view 005 learning, prioritizes the complementarity while 006 ignoring the consensus. To address this prob-007 lem, we propose an enhanced multi-head self-008 attention (EMHA). First, to satisfy the comple-009 mentary principle, EMHA removes the one-010 to-one mapping constraint among queries and 011 keys in multiple subspaces and allows each 012 query to attend to multiple keys. On top of that, 013 we develop a method to fully encourage consen-014 sus among heads by introducing two interaction 015 models, namely Inner-Subspace Interaction and 016 Cross-Subspace Interaction. Extensive experi-017 ments on a wide range of language tasks (e.g., 018 machine translation, abstractive summarization 019 and grammar correction, language modeling),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Brain Network TransformerXuan Kan, Wei Dai, Hejie Cui, Zilong Zhang 等NeurIPS 2022 · 被引用 272 次
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 被引用 212 次
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 被引用 113 次
- Revisiting Over-smoothing in BERT from the Perspective of GraphHan Shi, Jiahui Gao, Hang Xu, Xiaodan Liang 等ICLR 2022 · 被引用 92 次
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao 等ICML 2022 · 被引用 15 次
相关 Paper
- Enlivening Redundant Heads in Multi-head Self-attention for Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoEMNLP 2021 · 被引用 9 次
- Contrastive Multi-view Subspace Clustering via Tensor Transformers AutoencoderQianqian Wang, Zihao Zhang, Wei Feng, Zhiqiang Tao 等AAAI 2025 · 被引用 5 次
- Finding the Pillars of Strength for Multi-Head AttentionJinjie Ni, Rui Mao, Zonglin Yang, Han Lei 等ACL 2023 · 被引用 9 次
- Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsHuimin Huang, Yawen Huang, Lanfen Lin, Ruofeng Tong 等CVPR 2024
- Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningSiming Lan, Rui Zhang, Qi Yi, Jiaming Guo 等NeurIPS 2023 · 被引用 18 次
