CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
Jingyu Lei, Gaoang Wang, Der-Horng Lee
Abstract
Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a high-level semantic understanding, leading to suboptimal merges, information redundancy, or context loss. To address these limitations, we introduce CORE (Compact Object-centric REpresentations), a new paradigm for visual token compression. CORE leverages an efficient segmentation decoder to generate object masks, which serve as a high-level semantic prior to guide the merging of visual tokens into a compact set of object-centric representations. Furthermore, a novel centroid-guided sorting mechanism restores a coherent spatial order to the merged tokens, preserving vital positional information. Extensive experiments show that CORE not only establishes a new state-of-the-art on six authoritative benchmarks for fixed-rate compression, but also achieves dramatic efficiency gains in adaptive-rate settings. Even under extreme compression, after aggressively retaining with only 2.2% of all visual tokens, CORE still maintains 97.4% of baseline performance. Our work demonstrates the superiority of object-centric representations for efficient and effective LVLM processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 236a10d7-957e-4abb-9b87-4fbf925793a6Builds on54
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language ModelsTong Xu, Hailong Shi, Xingyu GaoCVPR 2026
- One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMsYongru Chen, Kai Zhang, Zeliang Zong, Yuchen Lu et al.CVPR 2026 · 1 citation
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu et al.NeurIPS 2025 · 32 citations
- Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept CoverageXueting Li, Qi Liu, Chenghao Xu, Xu Yang et al.ACL 2026
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai et al.CVPR 2026 · 7 citations
