EIT: Enhanced Interactive Transformer
Tong Zheng, Bei Li, Huiwen Bao, Tong Xiao, JingBo Zhu
Abstract
Two principles: the complementary princi-001 ple and the consensus principle are widely 002 acknowledged in the literature of multi-view 003 learning. However, the current design of Multi-004 head self-attention, an instance of multi-view 005 learning, prioritizes the complementarity while 006 ignoring the consensus. To address this prob-007 lem, we propose an enhanced multi-head self-008 attention (EMHA). First, to satisfy the comple-009 mentary principle, EMHA removes the one-010 to-one mapping constraint among queries and 011 keys in multiple subspaces and allows each 012 query to attend to multiple keys. On top of that, 013 we develop a method to fully encourage consen-014 sus among heads by introducing two interaction 015 models, namely Inner-Subspace Interaction and 016 Cross-Subspace Interaction. Extensive experi-017 ments on a wide range of language tasks (e.g., 018 machine translation, abstractive summarization 019 and grammar correction, language modeling),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78a64775-ccea-42b2-9a5c-348de6aec948Builds on6
- Brain Network TransformerXuan Kan, Wei Dai, Hejie Cui, Zilong Zhang et al.NeurIPS 2022 · 272 citations
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 212 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
- Revisiting Over-smoothing in BERT from the Perspective of GraphHan Shi, Jiahui Gao, Hang Xu, Xiaodan Liang et al.ICLR 2022 · 92 citations
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao et al.ICML 2022 · 15 citations
Related papers
- Enlivening Redundant Heads in Multi-head Self-attention for Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoEMNLP 2021 · 9 citations
- Contrastive Multi-view Subspace Clustering via Tensor Transformers AutoencoderQianqian Wang, Zihao Zhang, Wei Feng, Zhiqiang Tao et al.AAAI 2025 · 5 citations
- Finding the Pillars of Strength for Multi-Head AttentionJinjie Ni, Rui Mao, Zonglin Yang, Han Lei et al.ACL 2023 · 9 citations
- Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsHuimin Huang, Yawen Huang, Lanfen Lin, Ruofeng Tong et al.CVPR 2024
- Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningSiming Lan, Rui Zhang, Qi Yi, Jiaming Guo et al.NeurIPS 2023 · 18 citations
