MAML-en-LLM: Model Agnostic Meta-Training of LLMs for Improved In-Context Learning
Sanchit Sinha, Yuguang Yue, Victor Soto, Mayank Kulkarni, Jianhua Lu, Aidong Zhang
摘要
Adapting large language models (LLMs) to unseen tasks with incontext training samples without fine-tuning remains an important research problem. To learn a robust LLM that adapts well to unseen tasks, multiple meta-training approaches have been proposed such as MetaICL and MetaICT, which involve meta-training pre-trained LLMs on a wide variety of diverse tasks. These meta-training approaches essentially perform in-context multi-task fine-tuning and evaluate on a disjointed test set of tasks. Even though they achieve impressive performance, their goal is never to compute a truly general set of parameters. In this paper, we propose MAML-en-LLM, a novel method for meta-training LLMs, which can learn truly generalizable parameters that not only performs well on disjointed tasks but also adapts to unseen tasks. We see an average increase of 2% on unseen domains in the performance while a massive 4% improvement on adaptation performance. Furthermore, we demonstrate that MAML-en-LLM outperforms baselines in settings with limited amount of training data on both seen and unseen domains by an average of 2%. Finally, we discuss the effects of type of tasks, optimizers and task complexity, an avenue barely explored in metatraining literature. Exhaustive experiments across 7 task settings along with two data settings demonstrate that models trained with MAML-en-LLM outperform SOTA meta-training approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ContextNav: Towards Agentic Multimodal In-Context LearningHonghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang 等ICLR 2026 · 被引用 14 次
- ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task AdaptationYihua Shao, Xiaofeng Lin, Xinwei Long, Siyu Chen 等AAAI 2026 · 被引用 8 次
- InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task PlanningZheng Wang, Shu Xian Teo, Jun Jie Chew, Wei ShiSIGIR 2025 · 被引用 4 次
- On the Stability and Generalization of Meta-Learning: the Impact of Inner-LevelsWenjun Ding, Jingling Liu, Lixing Chen, Xiu Su 等NeurIPS 2025 · 被引用 2 次
- Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLMLuo Ji, Qi Qin, Ningyuan Xi, Teng Chen 等ICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- Noisy Channel Language Model Prompting for Few-Shot Text ClassificationSewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke ZettlemoyerACL 2022 · 被引用 237 次
- Meta-Learning without MemorizationMingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine 等ICLR 2020 · 被引用 201 次
相关 Paper
- Meta-learning via Language Model In-context TuningYanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis 等ACL 2022
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 被引用 5 次
- Boosting Natural Language Generation from Instructions with Meta-LearningBudhaditya Deb, Ahmed Hassan Awadallah, Guoqing ZhengEMNLP 2022 · 被引用 5 次
- Meta-in-context learning in large language modelsJulian Coda-Forno, Marcel Binz, Zeynep Akata, Matt M. Botvinick 等NeurIPS 2023 · 被引用 81 次
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 被引用 19 次
