Visual Textualization for Image Prompted Object Detection
Yongjian Wu, Yang Zhou, Jiya Saiyin, Bingzheng Wei, Yan Xu
摘要
We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are difficult to describe textually and nearly absent from their pre-training data, while preserving their pre-trained object-text alignment. Specifically, VisTex-OVLM leverages multi-scale textualizing blocks and a multi-stage fusion strategy to integrate visual information from visual exemplars, generating textualized visual tokens that effectively guide OVLMs alongside text prompts. Unlike previous methods, our method maintains the original architecture of OVLM, maintaining its generalization capabilities while enhancing performance in few-shot settings. VisTex-OVLM demonstrates superior performance across open-set datasets which have minimal overlap with OVLM's pre-training data and achieves state-of-the-art results on few-shot benchmarks PASCAL VOC and MSCOCO. The code will be released at https://github.com/WitGotFlg/VisTex-OVLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object DetectionJiazhou Zhou, Qing Jiang, Kanghao Chen, Lutao Jiang 等AAAI 2026
- GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric RectificationJiadong Yan, Ke Zhang, Chenyang Zhao, Shoushan Li 等ICML 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
相关 Paper
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji 等ICCV 2025 · 被引用 1 次
- Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningKun Ding, Haojian Zhang, Qiang Yu, Ying Wang 等AAAI 2024 · 被引用 8 次
- OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesZehong Ma, Shiliang Zhang, Longhui Wei, Qi TianCVPR 2024 · 被引用 6 次
- LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained DescriptorsSheng Jin, Xueying Jiang, Jiaxing Huang, Lewei Lu 等ICLR 2024 · 被引用 48 次
- Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model BlindnessXin Hu, Haomiao Ni, Yunbei Zhang, Jihun Hamm 等CVPR 2026 · 被引用 1 次
