Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
Emily Xiao, Chin-Jou Li, Yilin Zhang, Graham Neubig, Amanda Bertsch
Abstract
Many-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks. However, this shifts the computational burden from training-time to inference-time, making deployment of many-shot ICL challenging to justify in-practice. This cost is further increased if a custom demonstration set is retrieved for each inference example. We present Dynamic Block-Sparse Attention, a training-free framework for retrieval-based many-shot in-context learning. By combining carefully designed blocksparse attention and retrieval of cached groups of demonstrations, we achieve comparable perexample latency to finetuning while maintaining on average >95% of the best method's accuracy across strong ICL and finetuning baselines. We hope that this will further enable the deployment of many-shot ICL at scale. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Prompt-MII: Meta-Learning Instruction Induction for LLMsEmily Xiao, Yixiao Zeng, Ada Chen, Chin-Jou Li et al.ICLR 2026 · 9 citations
- AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache ReuseJie Ou, Jinyu Guo, Shiyao Guo, Yuang Li et al.ACL 2026
Builds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- Focused Large Language Models are Stable Many-Shot LearnersPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.EMNLP 2024
- HiFICL: High-Fidelity In-Context Learning for Multimodal TasksXiaoyu Li, Yuhang Liu, xuanshuo kang, zheng luo et al.CVPR 2026 · 1 citation
- GistScore: Learning Better Representations for In-Context Example Selection with Gist BottlenecksShivanshu Gupta, Clemens Rosenbaum, Ethan R. ElenbergICML 2024 · 10 citations
- Unified Demonstration Retriever for In-Context LearningXiaonan Li, Kai Lv, Hang Yan, Tianyang Lin et al.ACL 2023 · 40 citations
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 3 citations
