Pre-Training to Learn in Context
Yuxian Gu, Li Dong, Furu Wei, Minlie Huang
Abstract
In-context learning, where pre-trained language models learn to perform tasks from task examples and instructions in their contexts, has attracted much attention in the NLP community. However, the ability of in-context learning is not fully exploited because language models are not explicitly trained to learn in context. To this end, we propose PICL (Pre-training for In-Context Learning), a framework to enhance the language models' in-context learning ability by pre-training the model on a large collection of "intrinsic tasks" in the general plain-text corpus using the simple language modeling objective. PICL encourages the model to infer and perform tasks by conditioning on the contexts while maintaining task generalization of pre-trained models. We evaluate the in-context learning performance of the model trained with PICL on seven widely-used text classification datasets and the Super-NaturalInstrctions benchmark, which contains 100+ NLP tasks formulated to text generation. Our experiments show that PICL is more effective and task-generalizable than a range of baselines, outperforming larger language models with nearly 4x parameters. The code is publicly available at https://github.com/thu-coai/PICL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 146 citations
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou et al.ICLR 2024 · 87 citations
- Supervised Knowledge Makes Large Language Models Better In-context LearnersLinyi Yang, Shuibai Zhang, Zhuohao Yu, Guangsheng Bao et al.ICLR 2024 · 28 citations
- On the Noise Robustness of In-Context Learning for Text GenerationHongfu Gao, Feipeng Zhang, Wenyu Jiang, Jun Shu et al.NeurIPS 2024 · 20 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 5 citations
- The Learnability of In-Context LearningNoam Wies, Yoav Levine, Amnon ShashuaNeurIPS 2023 · 207 citations
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 48 citations
- Understanding In-Context Learning via Supportive Pretraining DataXiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov et al.ACL 2023 · 16 citations
- Large-scale Lifelong Learning of In-context Instructions and How to Tackle ItJisoo Mok, Jaeyoung Do, Sungjin Lee, Tara Taghavi et al.ACL 2023 · 3 citations
