Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs
Shiang-Feng Tsai, Yuan-Hong Liao, Jin-Cheng Jhang, Nan Qiao, Min Sun
Abstract
Part-level pointing is important for fine-grained interaction and reasoning, yet existing Multimodal Large Language Models (MLLMs) remain limited to instance-level pointing. Part-level pointing is inherently more challenging than instance-level pointing, due to the fine-grained and ambiguous nature of object parts. Instead of relying on post-training, we explore a few-shot mechanism to enhance the pointing ability of MLLMs. We introduce POinting at Parts (POP), a training-free, plug-and-play approach that addresses the challenges of part-level pointing. POP fuses textual and visual attention maps with self-supervised visual correspondences from query image and few-shot examples. On average across the three evaluated datasets, POP achieves accuracy gains of up to 8.9 points in the one-shot setting and 16.4 points in the three-shot setting for the pointing-capable MLLMs-Qwen2.5-VL, Ovis2.5, and Molmo. Notably, even MLLMs without pointing capability benefit significantly from the proposed approach. These results establish a simple yet effective path toward fine-grained spatial grounding in MLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c887aac-4e49-4850-88be-fe72285c0d86Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 494 citations
Related papers
- Enhancing Part-Level Point Grounding for Any Open-Source MLLMsJin-Cheng Jhang, Fu-En Wang, Xin Yang, Nan Qiao et al.CVPR 2026
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsMingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li et al.NeurIPS 2024 · 50 citations
- Spatial Preference Rewarding for MLLMs Spatial UnderstandingHan Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang et al.ICCV 2025 · 3 citations
- Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual PromptingGiacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca MoroAAAI 2026
- Grounding Everything in Tokens for Multimodal Large Language ModelsXiangxuan Ren, Zhongdao Wang, Liping Hou, Pin Tang et al.CVPR 2026 · 2 citations
