What's in Your Head? Emergent Behaviour in Multi-Task Transformer Models
Mor Geva, Uri Katz, Aviv Ben-Arie, Jonathan Berant
摘要
The primary paradigm for multi-task training in natural language processing is to represent the input with a shared pre-trained language model, and add a small, thin network (head) per task. Given an input, a target head is the head that is selected for outputting the final prediction. In this work, we examine the behaviour of non-target heads, that is, the output of heads when given input that belongs to a different task than the one they were trained for. We find that non-target heads exhibit emergent behaviour, which may either explain the target task, or generalize beyond their original task. For example, in a numerical reasoning task, a span extraction head extracts from the input the arguments to a computation that results in a number generated by a target generative head. In addition, a summarization head that is trained with a target question answering head, outputs query-based summaries when given a question and a context from which the answer is to be extracted. This emergent behaviour suggests that multi-task training leads to nontrivial extrapolation of skills, which can be harnessed for interpretability and generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
- ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li 等EMNLP 2022 · 被引用 13 次
- Peek Across: Improving Multi-Document Modeling via Cross-Document Question-AnsweringAvi Caciularu, Matthew E. Peters, Jacob Goldberger, Ido Dagan 等ACL 2023 · 被引用 8 次
- When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPJingwei Ni, Zhijing Jin, Qian Wang, Mrinmaya Sachan 等ACL 2023 · 被引用 2 次
它引用的顶会 Paper6
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen 等EMNLP 2021 · 被引用 176 次
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 被引用 76 次
- Joint Learning of Answer Selection and Answer Summary Generation in Community Question AnsweringYang Deng, Wai Lam, Yuexiang Xie, Daoyuan Chen 等AAAI 2020 · 被引用 65 次
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 被引用 12 次
相关 Paper
- Interpreting and Exploiting Functional Specialization in Multi-Head Attention under Multi-task LearningChong Li, Shaonan Wang, Yunhao Zhang, Jiajun Zhang 等EMNLP 2023 · 被引用 5 次
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 被引用 111 次
- Extrapolation by Association: Length Generalization Transfer In TransformersZiyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak 等NeurIPS 2025 · 被引用 13 次
- Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasksTianyu He, Darshil Doshi, Aritra Das, Andrey GromovNeurIPS 2024 · 被引用 52 次
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 被引用 2 次
