Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling
Hongyu Gong, Yun Tang, Juan Miguel Pino, Xian Li
摘要
Multi-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Multilingual and multi-domain learning are common scenarios for sequence modeling, where the key challenge is to maximize positive transfer and mitigate negative transfer across languages and domains. In this paper, we find that non-selective attention sharing is sub-optimal for achieving good generalization across all languages and domains. We further propose attention sharing strategies to facilitate parameter sharing and specialization in multilingual and multi-domain sequence modeling. Our approach automatically learns shared and specialized attention heads for different languages and domains to mitigate their interference. Evaluated in various tasks including speech recognition, text-to-text and speech-to-text translation, the proposed attention sharing strategies consistently bring gains to sequence models built upon multi-head attention. For speech-to-text translation, our approach yields an average of BLEU over language directions in multilingual setting and BLEU over domains in multi-domain setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Learning to Branch for Multi-Task LearningPengsheng Guo, Chen-Yu Lee, Daniel UlbrichtICML 2020 · 被引用 208 次
- Understanding and Improving Information Transfer in Multi-Task LearningSen Wu, Hongyang R. Zhang, Christopher RéICLR 2020 · 被引用 183 次
- Go From the General to the Particular: Multi-Domain Translation with Domain Transformation NetworksYong Wang, Longyue Wang, Shuming Shi, Victor O. K. Li 等AAAI 2020 · 被引用 30 次
相关 Paper
- Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain MixingHaoming Jiang, Chen Liang, Chong Wang, Tuo ZhaoACL 2020 · 被引用 26 次
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang 等ACL 2021
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang 等EMNLP 2023 · 被引用 5 次
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 被引用 25 次
- Adaptive Token-level Cross-lingual Feature Mixing for Multilingual Neural Machine TranslationJunpeng Liu, Kaiyu Huang, Jiuyi Li, Huan Liu 等EMNLP 2022 · 被引用 5 次
