The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders
Han He, Jinho D. Choi
Abstract
Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency while a question still remains whether or not it would perform as well on tasks that are distinct in nature. We first present MTL results on five NLP tasks, POS, NER, DEP, CON, and SRL, and depict its deficiency over single-task learning. We then conduct an extensive pruning analysis to show that a certain set of attention heads get claimed by most tasks during MTL, who interfere with one another to fine-tune those heads for their own objectives. Based on this finding, we propose the Stem Cell Hypothesis to reveal the existence of attention heads naturally talented for many tasks that cannot be jointly trained to create adequate embeddings for all of those tasks. Finally, we design novel parameter-free probes to justify our hypothesis and demonstrate how attention heads are transformed across the five tasks during MTL through label analysis. Performance % of Attention Heads Kept PS/S STL STL-SP STL-DP MTL-DP STL-SP STL-DP MTL-DP STL STL-DP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c95d054-6d71-4305-bec1-88d30f169bd9Cited by top-tier papers9
- Empirical Study of Zero-Shot NER with ChatGPTTingyu Xie, Qi Li, Jian Zhang, Yan Zhang et al.EMNLP 2023 · 50 citations
- A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt LearningJiangui Chen, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2023 · 31 citations
- Demystifying Privacy Policy of Third-Party Libraries in Mobile AppsKaifa Zhao, Xian Zhan, Le Yu, Shiyao Zhou et al.ICSE 2023 · 22 citations
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen et al.WWW 2025 · 16 citations
- ChinaOpen: A Dataset for Open-world Multimodal LearningAozhu Chen, Ziyuan Wang, Chengbo Dong, Kaibin Tian et al.ACM MM 2023 · 8 citations
Builds on4
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 35 citations
- Relabel the Noise: Joint Extraction of Entities and Relations via Cooperative MultiagentsDaoyuan Chen, Yaliang Li, Kai Lei, Ying ShenACL 2020 · 13 citations
Related papers
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang et al.EMNLP 2023 · 5 citations
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang et al.ACL 2021
- What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsMor Geva, Uri Katz, Aviv Ben-Arie, Jonathan BerantEMNLP 2021
- An Information-theoretic Multi-task Representation Learning Framework for Natural Language UnderstandingDou Hu, Lingwei Wei, Wei Zhou, Songlin HuAAAI 2025 · 3 citations
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 32 citations
