The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders
Han He, Jinho D. Choi
摘要
Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency while a question still remains whether or not it would perform as well on tasks that are distinct in nature. We first present MTL results on five NLP tasks, POS, NER, DEP, CON, and SRL, and depict its deficiency over single-task learning. We then conduct an extensive pruning analysis to show that a certain set of attention heads get claimed by most tasks during MTL, who interfere with one another to fine-tune those heads for their own objectives. Based on this finding, we propose the Stem Cell Hypothesis to reveal the existence of attention heads naturally talented for many tasks that cannot be jointly trained to create adequate embeddings for all of those tasks. Finally, we design novel parameter-free probes to justify our hypothesis and demonstrate how attention heads are transformed across the five tasks during MTL through label analysis. Performance % of Attention Heads Kept PS/S STL STL-SP STL-DP MTL-DP STL-SP STL-DP MTL-DP STL STL-DP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Empirical Study of Zero-Shot NER with ChatGPTTingyu Xie, Qi Li, Jian Zhang, Yan Zhang 等EMNLP 2023 · 被引用 50 次
- A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt LearningJiangui Chen, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等SIGIR 2023 · 被引用 31 次
- Demystifying Privacy Policy of Third-Party Libraries in Mobile AppsKaifa Zhao, Xian Zhan, Le Yu, Shiyao Zhou 等ICSE 2023 · 被引用 22 次
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen 等WWW 2025 · 被引用 16 次
- ChinaOpen: A Dataset for Open-world Multimodal LearningAozhu Chen, Ziyuan Wang, Chengbo Dong, Kaibin Tian 等ACM MM 2023 · 被引用 8 次
它引用的顶会 Paper4
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 被引用 35 次
- Relabel the Noise: Joint Extraction of Entities and Relations via Cooperative MultiagentsDaoyuan Chen, Yaliang Li, Kai Lei, Ying ShenACL 2020 · 被引用 13 次
相关 Paper
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang 等EMNLP 2023 · 被引用 5 次
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang 等ACL 2021
- What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsMor Geva, Uri Katz, Aviv Ben-Arie, Jonathan BerantEMNLP 2021
- An Information-theoretic Multi-task Representation Learning Framework for Natural Language UnderstandingDou Hu, Lingwei Wei, Wei Zhou, Songlin HuAAAI 2025 · 被引用 3 次
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 被引用 32 次
