Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning
Zhirui Chen, Ziwei Chen, Ling Shao
Abstract
Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by finite context windows and the growing cost of key–value (KV) caches in long multi-modal sequences. Existing memory compression approaches typically rely on rigid token removal or sample-dependent importance estimation, which introduces bias, disrupts semantic structure—particularly for visual representations—and yields static memories that cannot adapt to new queries. We introduce TASM (Task-Aware Structured Memory), a training-free framework that addresses these limitations through task-aware, structure-preserving, and dynamically accessible memory construction. TASM employs Task-Vector Guided Compression to replace sample-specific signals with a task-level direction that captures shared relevance across demonstrations. To preserve the underlying information manifold, it further applies Semantics-Aware Token Merging, formulating compression as a Bipartite Graph Matching problem that merges tokens without destructive pruning. Finally, TASM organizes compressed representations into a multi-resolution hierarchy consisting of a compact Core Memory and a Latent Bank, enabling Query-Adaptive Dynamic Activation and Dynamic Retrieval at inference time. Empirical evaluations show that TASM sustains strong multi-modal ICL performance under high compression ratios, demonstrating an effective balance between efficiency, adaptability, and semantic fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08dbc3cc-e45c-40f2-a9ec-3d6b903cfb59Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- Efficient Multi-modal Long Context Learning for Training-free AdaptationZehong Ma, Shiliang Zhang, Longhui Wei, Qi TianICML 2025
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 1 citation
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen et al.ICML 2026
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang et al.ICLR 2026 · 3 citations
- MMSep: Efficient Multimodal Long-Generation Reasoning via Multimodal Separator CompressionMingjie Ma, Yichao Ma, Jiannan Cao, Changhong Li et al.KDD 2026
