What's in Your Head? Emergent Behaviour in Multi-Task Transformer Models
Mor Geva, Uri Katz, Aviv Ben-Arie, Jonathan Berant
Abstract
The primary paradigm for multi-task training in natural language processing is to represent the input with a shared pre-trained language model, and add a small, thin network (head) per task. Given an input, a target head is the head that is selected for outputting the final prediction. In this work, we examine the behaviour of non-target heads, that is, the output of heads when given input that belongs to a different task than the one they were trained for. We find that non-target heads exhibit emergent behaviour, which may either explain the target task, or generalize beyond their original task. For example, in a numerical reasoning task, a span extraction head extracts from the input the arguments to a computation that results in a number generated by a target generative head. In addition, a summarization head that is trained with a target question answering head, outputs query-based summaries when given a question and a context from which the answer is to be extracted. This emergent behaviour suggests that multi-task training leads to nontrivial extrapolation of skills, which can be harnessed for interpretability and generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
- ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li et al.EMNLP 2022 · 13 citations
- Peek Across: Improving Multi-Document Modeling via Cross-Document Question-AnsweringAvi Caciularu, Matthew E. Peters, Jacob Goldberger, Ido Dagan et al.ACL 2023 · 8 citations
- When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPJingwei Ni, Zhijing Jin, Qian Wang, Mrinmaya Sachan et al.ACL 2023 · 2 citations
Builds on6
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen et al.EMNLP 2021 · 176 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- Joint Learning of Answer Selection and Answer Summary Generation in Community Question AnsweringYang Deng, Wai Lam, Yuexiang Xie, Daoyuan Chen et al.AAAI 2020 · 65 citations
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 12 citations
Related papers
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang et al.EMNLP 2023 · 5 citations
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 111 citations
- Extrapolation by Association: Length Generalization Transfer In TransformersZiyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak et al.NeurIPS 2025 · 13 citations
- Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasksTianyu He, Darshil Doshi, Aritra Das, Andrey GromovNeurIPS 2024 · 52 citations
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 2 citations
