Not All Tasks Are Born Equal: Understanding Zero-Shot Generalization
Jing Zhou, Zongyu Lin, Yanan Zheng, Jian Li, Zhilin Yang
Abstract
Recent work has achieved remarkable zero-shot performance with multi-task prompted pretraining, but little has been understood. For the first time, we show that training on a small number of key tasks beats using all the training tasks, while removing these key tasks substantially hurts performance. We also find that these key tasks are mostly question answering (QA) tasks. These novel findings combined deepen our understanding about zero-shot generalization—training on certain tasks such as QA encodes general knowledge transferable to a wide range of tasks. In addition, to automate this procedure, we devise a method that (1) identifies key training tasks without observing the test tasks by examining the pairwise generalization results and (2) resamples training tasks for better data distribution. Empirically, our approach achieves improved results across various model scales and tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers11
- Localizing Task Information for Improved Model Merging and CompressionKe Wang, Nikolaos Dimitriadis, Guillermo Ortiz-Jiménez, François Fleuret et al.ICML 2024 · 107 citations
- EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerceYangning Li, Shirong Ma, Xiaobin Wang, Shen Huang et al.AAAI 2024 · 85 citations
- Learning to Route Among Specialized Experts for Zero-Shot GeneralizationMohammed Muqeeth, Haokun Liu, Yufan Liu, Colin RaffelICML 2024 · 63 citations
- Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive TasksPo-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang et al.EMNLP 2023 · 7 citations
- Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific TasksChangho Lee, Janghoon Han, Seonghyeon Ye, Stanley Jungkyu Choi et al.EMNLP 2024 · 1 citation
Related papers
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 8 citations
- TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text ClassificationChengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang et al.EMNLP 2021 · 39 citations
- Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsKabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit ChoudhuryACL 2022
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu et al.ACL 2023 · 22 citations
