Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task Learning
Chong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang, Chengqing Zong
摘要
Transformer-based models, even though achieving super-human performance on several downstream tasks, are often regarded as a black box and used as a whole. It is still unclear what mechanisms they have learned, especially their core module: multi-head attention. Inspired by functional specialization in the human brain, which helps to efficiently handle multiple tasks, this work attempts to figure out whether the multi-head attention module will evolve similar function separation under multitasking training. If it is, can this mechanism further improve the model performance? To investigate these questions, we introduce an interpreting method to quantify the degree of functional specialization in multi-head attention. We further propose a simple multi-task training method to increase functional specialization and mitigate negative information transfer in multi-task learning. Experimental results on seven pre-trained transformer models have demonstrated that multi-head attention does evolve functional specialization phenomenon after multi-task training which is affected by the similarity of tasks. Moreover, the multi-task training strategy based on functional specialization boosts performance in both multi-task learning and transfer learning without adding any parameters. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Fine-Grained Activation Steering: Steering Less, Achieving MoreZijian Feng, Tianjiao Li, Zixiao Zhu, Hanzhang Zhou 等ICLR 2026 · 被引用 6 次
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao 等EMNLP 2024 · 被引用 4 次
- Reallocating Attention Across Layers to Reduce Multimodal HallucinationHaolang Lu, Bolun Chu, WeiYe Fu, Guoshun Nan 等CVPR 2026 · 被引用 3 次
- HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly DetectionZhaolin Cai, Fan Li, Ziwei Zheng, Haixia Bi 等AAAI 2026 · 被引用 1 次
- Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering VectorsWeixuan Wang, Jingyuan Yang, Wei PengICLR 2025
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 被引用 282 次
- Understanding and Improving Information Transfer in Multi-Task LearningSen Wu, Hongyang R. Zhang, Christopher RéICLR 2020 · 被引用 183 次
相关 Paper
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 被引用 111 次
- Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine TranslationShaomu Tan, Di Wu, Christof MonzEMNLP 2024
- What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsMor Geva, Uri Katz, Aviv Ben-Arie, Jonathan BerantEMNLP 2021
- Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingSenyu Han, Hongchuan Zeng, Kai Yu, Lu ChenICML 2025
- Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task LearningWooseong Jeong, Kuk-Jin YoonICCV 2025 · 被引用 2 次
