Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks
Anwoy Chatterjee, Eshaan Tanwar, Subhabrata Dutta, Tanmoy Chakraborty
摘要
Large Language Models (LLMs) have transformed NLP with their remarkable In-context Learning (ICL) capabilities. Automated assistants based on LLMs are gaining popularity; however, adapting them to novel tasks is still challenging. While colossal models excel in zero-shot performance, their computational demands limit widespread use, and smaller language models struggle without context. This paper investigates whether LLMs can generalize from labeled examples of predefined tasks to novel tasks. Drawing inspiration from biological neurons and the mechanistic interpretation of the Transformer architecture, we explore the potential for information sharing across tasks. We design a cross-task prompting setup with three LLMs and show that LLMs achieve significant performance improvements despite no examples from the target task in the context. Cross-task prompting leads to a remarkable performance boost of 107% for LLaMA-2 7B, 18.6% for LLaMA-2 13B, and 3.2% for GPT 3.5 on average over zeroshot prompting, and performs comparable to standard in-context learning. The effectiveness of generating pseudo-labels for in-task examples is demonstrated, and our analyses reveal a strong correlation between the effect of crosstask examples and model activation similarities in source and target input tokens. This paper offers a first-of-its-kind exploration of LLMs' ability to solve novel tasks based on contextual signals from different task examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text GenerationFrançois Ledoyen, Gaël Dias, Jérémie Pantin, Alexis Lechervy 等EMNLP 2025 · 被引用 1 次
- ExpeTrans: LLMs Are Experiential Transfer LearnersJinglong Gao, Xiao Ding, Lingxiao Zou, Bibo Cai 等ACL 2025
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- STraTA: Self-Training with Task Augmentation for Better Few-shot LearningTu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon 等EMNLP 2021 · 被引用 25 次
- Symbol tuning improves in-context learning in language modelsJerry W. Wei, Le Hou, Andrew K. Lampinen, Xiangning Chen 等EMNLP 2023 · 被引用 25 次
- Multilingual LLMs are Better Cross-lingual In-context Learners with AlignmentEshaan Tanwar, Subhabrata Dutta, Manish Borthakur, Tanmoy ChakrabortyACL 2023 · 被引用 21 次
相关 Paper
- Universal Self-Adaptive PromptingXingchen Wan, Ruoxi Sun, Hootan Nakhost, Hanjun Dai 等EMNLP 2023 · 被引用 4 次
- Enhancing LLM's Cognition via StructurizationKai Liu, Zhihang Fu, Chao Chen, Wei Zhang 等NeurIPS 2024 · 被引用 3 次
- Agent Instructs Large Language Models to be General Zero-Shot ReasonersNicholas Crispino, Kyle Montgomery, Fankun Zeng, Dawn Song 等ICML 2024 · 被引用 41 次
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee 等ICLR 2024 · 被引用 131 次
- Instance-adaptive Zero-shot Chain-of-Thought PromptingXiaosong Yuan, Chen Shen, Shaotian Yan, Xiaofeng Zhang 等NeurIPS 2024 · 被引用 46 次
