Lune

ACL2024Top-tier venue

EIT: Enhanced Interactive Transformer

Tong Zheng, Bei Li, Huiwen Bao, Tong Xiao, JingBo Zhu

2024Year

Abstract

Two principles: the complementary princi-001 ple and the consensus principle are widely 002 acknowledged in the literature of multi-view 003 learning. However, the current design of Multi-004 head self-attention, an instance of multi-view 005 learning, prioritizes the complementarity while 006 ignoring the consensus. To address this prob-007 lem, we propose an enhanced multi-head self-008 attention (EMHA). First, to satisfy the comple-009 mentary principle, EMHA removes the one-010 to-one mapping constraint among queries and 011 keys in multiple subspaces and allows each 012 query to attend to multiple keys. On top of that, 013 we develop a method to fully encourage consen-014 sus among heads by introducing two interaction 015 models, namely Inner-Subspace Interaction and 016 Cross-Subspace Interaction. Extensive experi-017 ments on a wide range of language tasks (e.g., 018 machine translation, abstractive summarization 019 and grammar correction, language modeling),

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 78a64775-ccea-42b2-9a5c-348de6aec948

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines