Lune

AAAI2026顶会

LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit

Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu, Yumeng Shi, Jinyang Guo, Wenya Wang

2026年份
3被引次数
3顶会引用

摘要

Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Our code is available at https://github.com/ModelTC/LightCompress . Recently, Large Language Models (LMMs) (Touvron et al. 2023; Liu et al. 2024a; Brown et al. 2020 ) have achieved rapid advancements in Natural Language Processing (NLP), which has become a significant milestone in the AI revolution. This breakthrough has quickly extended to vision modalities: mainstream Vision Language Models (VLMs) (Liu et al. 2023 (Liu et al. , 2024b;; Wang et al. 2024a; Chen et al. 2024c) typically encode visual inputs into tokens and unify multiple modalities within a shared embedding space, demonstrating strong visual-language understanding and generation capabilities in various tasks (Singh et al. 2019; Antol et al. 2015; Hudson and Manning 2019) .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper37

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖