Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng, Yiyu Wang, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou, Yuqian Fu, Bin Ren, Linfeng Zhang, Xuming Hu
摘要
Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly evaluated by measuring the accuracy drop on existing MLLM benchmarks before and after compression. However, these benchmarks are originally designed to assess general perception and reasoning abilities, rather than the specific challenges posed by visual token compression, leading to a fundamental task mismatch. In this work, we uncover a counterintuitive yet consistent phenomenon: simple image downsampling outperforms many advanced visual token compression methods across multiple widely used benchmarks. Through a comprehensive empirical study spanning eight popular benchmarks and multiple state-of-theart compression techniques, we show that (i) current benchmarks contain substantial noise (task-irrelevant samples) for evaluating visual token compression, and (ii) downsampling can act as an effective data filter that distinguishes between simple and difficult samples with respect to compression sensitivity. Motivated by these findings, we propose VTC-Bench 12 , an evaluation framework that explicitly leverages downsampling as a discriminator to denoise existing benchmarks, enabling a fairer and more meaningful additional assessment of visual token compression methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language ModelsYingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu ShenCVPR 2026 · 被引用 6 次
- VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactionsAdrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Yassine Ouali 等CVPR 2026
它引用的顶会 Paper12
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao 等AAAI 2024 · 被引用 314 次
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee 等ICCV 2025 · 被引用 37 次
- What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of GraphYutao Jiang, Qiong Wu, Wenhao Lin, Wei Yu 等AAAI 2025 · 被引用 27 次
相关 Paper
- Variation-aware Vision Token Dropping for Faster Large Vision-Language ModelsChen junjie, Xuyang Liu, Zichen Wen, Yiyu Wang 等CVPR 2026 · 被引用 24 次
- EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language ModelsZekun Wang, Minghua Ma, Zexin Wang, Rongchuan Mu 等ACL 2025
- Inference Optimal VLMs Need Fewer Visual Tokens and More ParametersKevin Y. Li, Sachin Goyal, João D. Semedo, J. Zico KolterICLR 2025
- Benchmarking and Enhancing VLM for Compressed Image UnderstandingZifu Zhang, Tongda Xu, Siqi Li, Shengxi Li 等ICML 2026 · 被引用 2 次
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang 等ICLR 2026 · 被引用 3 次
