Teaching VLMs to Localize Specific Objects from In-Context Examples
Sivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne, Raja Giryes, Rogério Feris, Leonid Karlinsky, James R. Glass, Assaf Arbelle, Shimon Ullman, Muhammad Jehanzeb Mirza
Abstract
Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efb80355-8d10-4e7d-a153-f17253ffdf56Cited by top-tier papers4
- Find your Needle: Small Object Image Retrieval via Multi-Object Attention OptimizationMichael Green, Matan Levy, Issar Tzachor, Dvir Samuel et al.NeurIPS 2025 · 1 citation
- Contextualized Visual Personalization in Vision-Language ModelsYeongtak Oh, Sangwon Yu, Junsung Park, Han Cheol Moon et al.ICML 2026 · 1 citation
- FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy OptimizationMohammed Asad Karim, Vinay Kumar VermaICML 2026 · 1 citation
- Can We Infer Object Pose Changes from Hand Movements?Julien Berry, Emmanuel Pietriga, Olivier Chapuis, Caroline AppertCHI 2026 · 1 citation
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- PIN: Positional Insert Unlocks Object Localisation Abilities in VLMsMichael Dorkenwald, Nimrod Barazani, Cees G. M. Snoek, Yuki M. AsanoCVPR 2024
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsKanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo et al.CVPR 2024 · 21 citations
- Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution DetectionFanhu Zeng, Zhen Cheng, Fei Zhu, Hongxin Wei et al.ICLR 2025
- DetGPT: Detect What You Need via ReasoningRenjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan et al.EMNLP 2023 · 57 citations
- CALVIN: Improved Contextual Video Captioning via Instruction TuningGowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri et al.NeurIPS 2024 · 4 citations
