Should We Be Pre-training? An Argument for End-task Aware Training as an Alternative
Lucio M. Dery, Paul Michel, Ameet Talwalkar, Graham Neubig
Abstract
In most settings of practical concern, machine learning practitioners know in advance what end-task they wish to boost with auxiliary tasks. However, widely used methods for leveraging auxiliary data like pre-training and its continued-pretraining variant are end-task agnostic: they rarely, if ever, exploit knowledge of the target task. We study replacing end-task agnostic continued training of pre-trained language models with end-task aware training of said models. We argue that for sufficiently important end-tasks, the benefits of leveraging auxiliary data in a task-aware fashion can justify forgoing the traditional approach of obtaining generic, end-task agnostic representations as with (continued) pre-training. On three different low-resource NLP tasks from two domains, we demonstrate that multi-tasking the end-task and auxiliary objectives results in significantly better downstream task performance than the widely-used task-agnostic continued pre-training paradigm of Gururangan et al. (2020). We next introduce an online meta-learning algorithm that learns a set of multi-task weights to better balance among our multiple auxiliary objectives, achieving further improvements on end-task performance and data efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32527972-53ef-4bf9-bbf3-b63ef3df63efCited by top-tier papers5
- Learning to Scaffold: Optimizing Model Explanations for TeachingPatrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins et al.NeurIPS 2022 · 26 citations
- Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary DataAlon Albalak, Colin A. Raffel, William Yang WangNeurIPS 2023 · 17 citations
- Instruction-tuned Language Models are Better Knowledge LearnersZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodríguez et al.ACL 2024 · 11 citations
- Continuous pseudo-labeling from the startDan Berrebbi, Ronan Collobert, Samy Bengio, Navdeep Jaitly et al.ICLR 2023 · 6 citations
- DoMIX: An Efficient Framework for Exploiting Domain Knowledge in Fine-TuningDohoon Kim, Donghun Kang, Taesup MoonACL 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
Related papers
- Learn to Cross-lingual Transfer with Meta Graph Learning Across Heterogeneous LanguagesZheng Li, Mukul Kumar, William Headden, Bing Yin et al.EMNLP 2020 · 26 citations
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang et al.EMNLP 2021 · 3 citations
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 19 citations
- Adaptive Transfer Learning on Graph Neural NetworksXueting Han, Zhenhuan Huang, Bang An, Jing BaiKDD 2021 · 30 citations
- TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text ClassificationChengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang et al.EMNLP 2021 · 39 citations
