STraTA: Self-Training with Task Augmentation for Better Few-shot Learning
Tu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon, Mohit Iyyer
摘要
Despite their recent successes in tackling many NLP tasks, large-scale pre-trained language models do not perform as well in few-shot settings where only a handful of training examples are available. To address this shortcoming, we propose STRATA, which stands for Self-Training with Task Augmentation, an approach that builds on two key ideas for effective leverage of unlabeled data. First, STRATA uses task augmentation, a novel technique that synthesizes a large amount of data for auxiliary-task fine-tuning from target-task unlabeled texts. Second, STRATA performs selftraining by further fine-tuning the strong base model created by task augmentation on a broad distribution of pseudo-labeled data. Our experiments demonstrate that STRATA can substantially improve sample efficiency across 12 fewshot benchmarks. Remarkably, on the SST-2 sentiment dataset, STRATA, with only 8 training examples per class, achieves comparable results to standard fine-tuning with 67K training examples. Our analyses reveal that task augmentation and self-training are both complementary and independently effective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' 等ACL 2022 · 被引用 332 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
- Towards Few-Shot Adaptation of Foundation Models via Multitask FinetuningZhuoyan Xu, Zhenmei Shi, Junyi Wei, Fangzhou Mu 等ICLR 2024 · 被引用 39 次
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi 等ACM MM 2023 · 被引用 20 次
- Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual GenerationTu Vu, Aditya Barua, Brian Lester, Daniel Cer 等EMNLP 2022 · 被引用 18 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong BaselinesMarius Mosbach, Maksym Andriushchenko, Dietrich KlakowICLR 2021 · 被引用 448 次
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' 等ACL 2022 · 被引用 332 次
相关 Paper
- Revisiting Self-training for Few-shot Learning of Language ModelYiming Chen, Yan Zhang, Chen Zhang, Grandee Lee 等EMNLP 2021 · 被引用 35 次
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang 等ICML 2023 · 被引用 64 次
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai 等EMNLP 2021 · 被引用 18 次
- Self-Supervised Meta-Learning for Few-Shot Natural Language Classification TasksTrapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, Andrew McCallumEMNLP 2020 · 被引用 9 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
