Triplet Attention: Rethinking the Similarity in Transformers
Haoyi Zhou, Jianxin Li, Jieqi Peng, Shuai Zhang, Shanghang Zhang
摘要
The Transformer model has benefited various real-world applications, where the self-attention mechanism with dot-products shows superior alignment ability on building long dependency. However, the pair-wisely attended self-attention limits further performance improvement on challenging tasks. To the extent of our knowledge, this is the first work to define the Triplet Attention (A3) for Transformer, which introduces triplet connections as the complementary dependency. Specifically, we define the triplet attention based on the scalar triplet product, which may be interchangeably used with the canonical one within the multi-head attention. It allows the self-attention mechanism to attend to diverse triplets and capture complex dependency. Then, we utilize the permuted formulation and kernel tricks to establish a linear approximation to A3. The proposed architecture could be smoothly integrated into the pre-training by modifying head configurations. Extensive experiments show that our methods achieve significant performance improvement on various tasks and two benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng 等ACL 2026
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 被引用 18 次
- How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker ComputationJosh Alman, Zhao SongICLR 2024 · 被引用 53 次
- Representational Strengths and Limitations of TransformersClayton Sanford, Daniel J. Hsu, Matus TelgarskyNeurIPS 2023 · 被引用 162 次
- Compositional Attention: Disentangling Search and RetrievalSarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio 等ICLR 2022 · 被引用 20 次
