Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
Emily Xiao, Chin-Jou Li, Yilin Zhang, Graham Neubig, Amanda Bertsch
摘要
Many-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks. However, this shifts the computational burden from training-time to inference-time, making deployment of many-shot ICL challenging to justify in-practice. This cost is further increased if a custom demonstration set is retrieved for each inference example. We present Dynamic Block-Sparse Attention, a training-free framework for retrieval-based many-shot in-context learning. By combining carefully designed blocksparse attention and retrieval of cached groups of demonstrations, we achieve comparable perexample latency to finetuning while maintaining on average >95% of the best method's accuracy across strong ICL and finetuning baselines. We hope that this will further enable the deployment of many-shot ICL at scale. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prompt-MII: Meta-Learning Instruction Induction for LLMsEmily Xiao, Yixiao Zeng, Ada Chen, Chin-Jou Li 等ICLR 2026 · 被引用 9 次
- AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache ReuseJie Ou, Jinyu Guo, Shiyao Guo, Yuang Li 等ACL 2026
它引用的顶会 Paper13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- Focused Large Language Models are Stable Many-Shot LearnersPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang 等EMNLP 2024
- HiFICL: High-Fidelity In-Context Learning for Multimodal TasksXiaoyu Li, Yuhang Liu, xuanshuo kang, zheng luo 等CVPR 2026 · 被引用 1 次
- GistScore: Learning Better Representations for In-Context Example Selection with Gist BottlenecksShivanshu Gupta, Clemens Rosenbaum, Ethan R. ElenbergICML 2024 · 被引用 10 次
- Unified Demonstration Retriever for In-Context LearningXiaonan Li, Kai Lv, Hang Yan, Tianyang Lin 等ACL 2023 · 被引用 40 次
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 被引用 3 次
