VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
Mingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li, Yue Qiu, Zekang Du, Mengyang Wu, Pingping Zhang, Kun Li, Hongzheng Yang, Wenao Ma, Jiaheng Wei
摘要
Multimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding boxes to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap raises uncertainty about whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and utilize them to solve problems. To address this limitation, we introduce VP-Bench, aiming to assess MLLMs' capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models' ability to perceive VPs in natural scenes, utilizing 30K visualized prompts spanning 8 shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 28 MLLMs, including proprietary systems (e.g., and open-source models (e.g., In-ternVL3 and Qwen2.5-VL). In addition, we provide a comprehensive analysis of factors affecting VP understanding, such as variations in VP attributes, question arrangement, and model scale. VP-Bench establishes a new reference framework for studying MLLMs' ability to comprehend and resolve grounded referring questions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing ModelsJiacheng Hou, Yining Sun, Ruochong Jin, Haochen Han 等ICML 2026 · 被引用 3 次
- MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMsTang Da Huang, Weidong Tang, Wen Qi Xu, Xianpeng GuoACL 2026
它引用的顶会 Paper15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
相关 Paper
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu 等ICLR 2026 · 被引用 8 次
- Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You WantWeifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao 等ICLR 2025
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 被引用 7 次
- ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual PromptsMu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer 等CVPR 2024
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsDavid Ma, Yuanxing Zhang, Jincheng Ren, Jiawei Guo 等ICLR 2026 · 被引用 5 次
