FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context Learning
Qinyuan Ye, Iz Beltagy, Matthew E. Peters, Xiang Ren, Hannaneh Hajishirzi
Abstract
Large pre-trained models are capable of few-shot in-context learning (ICL), i.e., performing a new task by prepending a few demonstrations before the test input. However, the concatenated demonstrations are often excessively long and induce additional computation. Inspired by fusion-in-decoder (FiD) models which efficiently aggregate more passages and thus outperforms concatenation-based models in open-domain QA, we hypothesize that similar techniques can be applied to improve the efficiency and end-task performance of ICL. To verify this, we present a comprehensive study on applying three fusion methods—concatenation-based (early fusion), FiD (intermediate), and ensemble-based (late)—to ICL. We adopt a meta-learning setup where a model is first trained to perform ICL on a mixture of tasks using one selected fusion method, then evaluated on held-out tasks for ICL. Results on 11 held-out tasks show that FiD-ICL matches or outperforms the other two fusion methods. Additionally, we show that FiD-ICL (1) is 10x faster at inference time compared to concat-based and ensemble-based ICL, as we can easily pre-compute the representations of in-context examples and reuse them; (2) enables scaling up to meta-training 3B-sized models, which would fail for concat-based ICL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebfab154-6f34-4f75-a0a5-a24c6de8fbddCited by top-tier papers6
- Supervised Knowledge Makes Large Language Models Better In-context LearnersLinyi Yang, Shuibai Zhang, Zhuohao Yu, Guangsheng Bao et al.ICLR 2024 · 28 citations
- Superposition Prompting: Improving and Accelerating Retrieval-Augmented GenerationThomas Merth, Qichen Fu, Mohammad Rastegari, Mahyar NajibiICML 2024 · 14 citations
- From Instance Training to Instruction Learning: Task Adapters Generation from InstructionsHuanxuan Liao, Shizhu He, Yao Xu, Yuanzhe Zhang et al.NeurIPS 2024 · 6 citations
- KBLaM: Knowledge Base augmented Language ModelXi Wang, Taketomo Isazawa, Liana Mikaelyan, James HensmanICLR 2025
- Mixtures of In-Context LearnersGiwon Hong, Emile van Krieken, Edoardo Maria Ponti, Nikolay Malkin et al.ACL 2025
Builds on15
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
Related papers
- What Do Language Models Learn in Context? The Structured Task HypothesisJiaoda Li, Yifan Hou, Mrinmaya Sachan, Ryan CotterellACL 2024 · 5 citations
- Optimizing Retrieval-augmented Reader Models via Token EliminationMoshe Berchansky, Peter Izsak, Avi Caciularu, Ido Dagan et al.EMNLP 2023 · 4 citations
- Revisiting In-context Learning Inference Circuit in Large Language ModelsHakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya InoueICLR 2025
- Large Language Models are Demonstration Pre-Selectors for ThemselvesJiarui Jin, Yuwei Wu, Haoxuan Li, Xiaoting He et al.ICML 2025
- HiFICL: High-Fidelity In-Context Learning for Multimodal TasksXiaoyu Li, Yuhang Liu, xuanshuo kang, zheng luo et al.CVPR 2026 · 1 citation
