STraTA: Self-Training with Task Augmentation for Better Few-shot Learning
Tu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon, Mohit Iyyer
Abstract
Despite their recent successes in tackling many NLP tasks, large-scale pre-trained language models do not perform as well in few-shot settings where only a handful of training examples are available. To address this shortcoming, we propose STRATA, which stands for Self-Training with Task Augmentation, an approach that builds on two key ideas for effective leverage of unlabeled data. First, STRATA uses task augmentation, a novel technique that synthesizes a large amount of data for auxiliary-task fine-tuning from target-task unlabeled texts. Second, STRATA performs selftraining by further fine-tuning the strong base model created by task augmentation on a broad distribution of pseudo-labeled data. Our experiments demonstrate that STRATA can substantially improve sample efficiency across 12 fewshot benchmarks. Remarkably, on the SST-2 sentiment dataset, STRATA, with only 8 training examples per class, achieves comparable results to standard fine-tuning with 67K training examples. Our analyses reveal that task augmentation and self-training are both complementary and independently effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 154ad8e9-d4fa-4abc-aef4-5acf3eff194eCited by top-tier papers13
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' et al.ACL 2022 · 332 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
- Towards Few-Shot Adaptation of Foundation Models via Multitask FinetuningZhuoyan Xu, Zhenmei Shi, Junyi Wei, Fangzhou Mu et al.ICLR 2024 · 39 citations
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
- Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual GenerationTu Vu, Aditya Barua, Brian Lester, Daniel Cer et al.EMNLP 2022 · 18 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 448 citations
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' et al.ACL 2022 · 332 citations
Related papers
- Revisiting Self-training for Few-shot Learning of Language ModelYiming Chen, Yan Zhang, Chen Zhang, Grandee Lee et al.EMNLP 2021 · 35 citations
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang et al.ICML 2023 · 64 citations
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai et al.EMNLP 2021 · 18 citations
- Self-Supervised Meta-Learning for Few-Shot Natural Language Classification TasksTrapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, Andrew McCallumEMNLP 2020 · 9 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
