NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, Zhilin Yang
摘要
Pretrained language models have become the standard approach for many NLP tasks due to strong performance, but they are very expensive to train. We propose a simple and efficient learning framework TLM that does not rely on large-scale pretraining 1 . Given some labeled task data and a large general corpus, TLM uses task data as queries to retrieve a tiny subset of the general corpus and jointly optimizes the task objective and the language modeling objective from scratch. On eight classification datasets in four domains, TLM achieves results better than or similar to pretrained language models (e.g., RoBERTa-Large) while reducing the training FLOPs by two orders of magnitude. With high accuracy and efficiency, we hope TLM will contribute to democratizing NLP and expediting its development 2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsJean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini 等NeurIPS 2023 · 被引用 63 次
- u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled ModalityWei-Ning Hsu, Bowen ShiNeurIPS 2022 · 被引用 58 次
- Should We Be Pre-training? An Argument for End-task Aware Training as an AlternativeLucio M. Dery, Paul Michel, Ameet Talwalkar, Graham NeubigICLR 2022 · 被引用 39 次
- Too Large; Data Reduction for Vision-Language Pre-TrainingAlex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei 等ICCV 2023 · 被引用 35 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
- UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset GenerationJuhwan Choi, Yeonghwa Kim, Seunguk Yu, Jungmin Yun 等EMNLP 2024 · 被引用 7 次
- Gradient-Free Structured Pruning with Unlabeled DataAzade Nova, Hanjun Dai, Dale SchuurmansICML 2023 · 被引用 38 次
- Vision-Language Model Selection and Reuse for Downstream AdaptationHao-Zhe Tan, Zhi Zhou, Yufeng Li, Lan-Zhe GuoICML 2025
- Landmark-Guided Policy Optimization for Multi-Objective Language Model SelectionMarcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft 等ICML 2026
