Should We Be Pre-training? An Argument for End-task Aware Training as an Alternative
Lucio M. Dery, Paul Michel, Ameet Talwalkar, Graham Neubig
摘要
In most settings of practical concern, machine learning practitioners know in advance what end-task they wish to boost with auxiliary tasks. However, widely used methods for leveraging auxiliary data like pre-training and its continued-pretraining variant are end-task agnostic: they rarely, if ever, exploit knowledge of the target task. We study replacing end-task agnostic continued training of pre-trained language models with end-task aware training of said models. We argue that for sufficiently important end-tasks, the benefits of leveraging auxiliary data in a task-aware fashion can justify forgoing the traditional approach of obtaining generic, end-task agnostic representations as with (continued) pre-training. On three different low-resource NLP tasks from two domains, we demonstrate that multi-tasking the end-task and auxiliary objectives results in significantly better downstream task performance than the widely-used task-agnostic continued pre-training paradigm of Gururangan et al. (2020). We next introduce an online meta-learning algorithm that learns a set of multi-task weights to better balance among our multiple auxiliary objectives, achieving further improvements on end-task performance and data efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning to Scaffold: Optimizing Model Explanations for TeachingPatrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins 等NeurIPS 2022 · 被引用 26 次
- Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary DataAlon Albalak, Colin A. Raffel, William Yang WangNeurIPS 2023 · 被引用 17 次
- Instruction-tuned Language Models are Better Knowledge LearnersZhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodríguez 等ACL 2024 · 被引用 11 次
- Continuous pseudo-labeling from the startDan Berrebbi, Ronan Collobert, Samy Bengio, Navdeep Jaitly 等ICLR 2023 · 被引用 6 次
- DoMIX: An Efficient Framework for Exploiting Domain Knowledge in Fine-TuningDohoon Kim, Donghun Kang, Taesup MoonACL 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
相关 Paper
- Learn to Cross-lingual Transfer with Meta Graph Learning Across Heterogeneous LanguagesZheng Li, Mukul Kumar, William Headden, Bing Yin 等EMNLP 2020 · 被引用 26 次
- Meta Distant Transfer Learning for Pre-trained Language ModelsChengyu Wang, Haojie Pan, Minghui Qiu, Jun Huang 等EMNLP 2021 · 被引用 3 次
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 被引用 19 次
- Adaptive Transfer Learning on Graph Neural NetworksXueting Han, Zhenhuan Huang, Bang An, Jing BaiKDD 2021 · 被引用 30 次
- TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text ClassificationChengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang 等EMNLP 2021 · 被引用 39 次
