Rapid Word Learning Through Meta In-Context Learning
Wentao Wang, Guangyuan Jiang, Tal Linzen, Brenden M. Lake
摘要
Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts. Yet the abilities of current language models for fewshot word learning, and methods for improving these abilities, are underexplored. In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow). This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word. This training is repeated on many new words to develop a general word-learning ability. We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pretrained on orders of magnitude more data. Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples. These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks. Word: aardvark Study examples: Look there's an aardvark, it's like an anteater. See the aardvark has a long snout for eating bugs. That must be the aardvark's house. Generalization example: The aardvark is hungry, it wants some snacks. Word: ski Study examples: Susie learned to ski last winter. People ski on tall mountains where there's lots of snow. I saw Susie ski fast down the snowy mountain. Generalization example: He will ski past the pine trees. Sentences: You can go fast or slow, and there are fun turns. Some animals hibernate in winter. Let's go to grandma's house! We warmed up by the fire.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 被引用 106 次
- Meta-in-context learning in large language modelsJulian Coda-Forno, Marcel Binz, Zeynep Akata, Matt M. Botvinick 等NeurIPS 2023 · 被引用 81 次
- Frequency Effects on Syntactic Rule Learning in TransformersJason Wei, Dan Garrette, Tal Linzen, Ellie PavlickEMNLP 2021 · 被引用 40 次
- Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change AnalysisMario Giulianelli, Iris Luden, Raquel Fernández, Andrey KutuzovACL 2023 · 被引用 9 次
相关 Paper
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang 等ICML 2023 · 被引用 64 次
- Self-Supervised Meta-Learning for Few-Shot Natural Language Classification TasksTrapit Bansal, Rishikesh Jha, Tsendsuren Munkhdalai, Andrew McCallumEMNLP 2020 · 被引用 9 次
- Making Pre-trained Language Models Better Few-shot LearnersTianyu Gao, Adam Fisch, Danqi ChenACL 2021
- MAML-en-LLM: Model Agnostic Meta-Training of LLMs for Improved In-Context LearningSanchit Sinha, Yuguang Yue, Victor Soto, Mayank Kulkarni 等KDD 2024 · 被引用 10 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
