Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts
Zhiwei Lin, Yongtao Wang, Zhi Tang
Abstract
Existing perception models achieve great success by learning from large amounts of labeled data, but they still struggle with open-world scenarios. To alleviate this issue, researchers introduce open-set perception tasks to detect or segment unseen objects in the training set. However, these models require predefined object categories as inputs during inference, which are not available in real-world scenarios. Recently, researchers pose a new and more practical problem, i.e., open-ended object detection, which discovers unseen objects without any object categories as inputs. In this paper, we present VL-SAM, a training-free framework that combines the generalized object recognition model (i.e., Vision-Language Model) with the generalized object localization model (i.e., Segment-Anything Model), to address the open-ended object detection and segmentation task. Without additional training, we connect these two generalized models with attention maps as the prompts. Specifically, we design an attention map generation module by employing head aggregation and a regularized attention flow to aggregate and propagate attention maps across all heads and layers in VLM, yielding high-quality attention maps. Then, we iteratively sample positive and negative points from the attention maps with a prompt generation module and send the sampled points to SAM to segment corresponding objects. Experimental results on the long-tail instance segmentation dataset (LVIS) show that our method surpasses the previous open-ended method on the object detection task and can provide additional instance segmentation masks. Besides, VL-SAM achieves favorable performance on the corner case object detection dataset (CODA), demonstrating the effectiveness of VL-SAM in real-world applications. Moreover, VL-SAM exhibits good model generalization that can incorporate various VLMs and SAMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea649d5c-bec3-45e3-862a-5de7a92e2833Cited by top-tier papers7
- InstructSAM: A Training-free Framework for Instruction-Oriented Remote Sensing Object RecognitionYijie Zheng, Weijie Wu, Qingyun Li, Xuehui Wang et al.NeurIPS 2025 · 12 citations
- VL-SAM-V2: Open-World Object Detection with General and Specific Query FusionZhiwei Lin, Yongtao WangNeurIPS 2025 · 6 citations
- Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning SegmentationSu Ho Han, Jeongseok Hyun, Pilhyeon Lee, Minho Shim et al.ICLR 2026 · 2 citations
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingXiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei et al.ICCV 2025 · 1 citation
- LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAsZixuan Hu, Yongxian Wei, Li Shen, Chun Yuan et al.CVPR 2025
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- WOW-Seg: A Word-free Open World Segmentation ModelDanyang Li, Tianhao Wu, Bin Lin, Zhenyuan Chen et al.ICLR 2026 · 2 citations
- Towards Universal Perception through Language-Guided Open-World Object DetectionZihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long et al.ACM MM 2025 · 1 citation
- Segment Anything, Even OccludedWei-En Tai, Yu-Lin Shih, Cheng Sun, Yu-Chiang Frank Wang et al.CVPR 2025
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
- Generative Region-Language Pretraining for Open-Ended Object DetectionChuang Lin, Yi Jiang, Lizhen Qu, Zehuan Yuan et al.CVPR 2024
