Fine-Grained Retrieval Prompt Tuning
Shijie Wang, Jianlong Chang, Zhihui Wang, Haojie Li, Wanli Ouyang, Qi Tian
Abstract
Fine-grained object retrieval aims to learn discriminative representation to retrieve visually similar objects. However, existing top-performing works usually impose pairwise similarities on the semantic embedding spaces or design a localization sub-network to continually fine-tune the entire model in limited data scenarios, thus resulting in convergence to suboptimal solutions. In this paper, we develop Fine-grained Retrieval Prompt Tuning (FRPT), which steers a frozen pre-trained model to perform the fine-grained retrieval task from the perspectives of sample prompting and feature adaptation. Specifically, FRPT only needs to learn fewer parameters in the prompt and adaptation instead of fine-tuning the entire model, thus solving the issue of convergence to suboptimal solutions caused by fine-tuning the entire model. Technically, a discriminative perturbation prompt (DPP) is introduced and deemed as a sample prompting process, which amplifies and even exaggerates some discriminative elements contributing to category prediction via a content-aware inhomogeneous sampling operation. In this way, DPP can make the fine-grained retrieval task aided by the perturbation prompts close to the solved task during the original pre-training. Thereby, it preserves the generalization and discrimination of representation extracted from input samples. Besides, a category-specific awareness head is proposed and regarded as feature adaptation, which removes the species discrepancies in features extracted by the pre-trained model using category-guided instance normalization. And thus, it makes the optimized features only include the discrepancies among subcategories. Extensive experiments demonstrate that our FRPT with fewer learnable parameters achieves the state-of-the-art performance on three widely-used fine-grained datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b7ca417-3046-43ab-a4a1-c73aa8b9b826Cited by top-tier papers7
- DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesXin Jiang, Hao Tang, Rui Yan, Jinhui Tang et al.ACM MM 2024 · 18 citations
- Learning to Parameterize Visual Attributes for Open-set Fine-grained RetrievalShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.NeurIPS 2023 · 13 citations
- Adversarial Reconstruction Feedback for Robust Fine-Grained GeneralizationShijie Wang, Jian Shi, Haojie LiICCV 2025 · 2 citations
- Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text RetrievalYifan Wang, Tao Wang, Chenwei Tang, Caiyang Yu et al.ACM MM 2025 · 1 citation
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Graph-Propagation Based Correlation Learning for Weakly Supervised Fine-Grained Image ClassificationZhuhui Wang, Shijie Wang, Haojie Li, Zhi Dou et al.AAAI 2020 · 107 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
Related papers
- Fine-Grained Image Retrieval via Dual-Vision AdaptationXin Jiang, Meiqi Cao, Hao Tang, Fei Shen et al.AAAI 2026 · 1 citation
- LPT: Long-tailed Prompt Tuning for Image ClassificationBowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng ZuoICLR 2023 · 19 citations
- CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language MisalignmentMaoyuan Shao, Yutong Gao, Xinyang Huang, Lijuan Sun et al.CVPR 2026 · 1 citation
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 16 citations
- PPT: Pre-trained Prompt Tuning for Few-shot LearningYuxian Gu, Xu Han, Zhiyuan Liu, Minlie HuangACL 2022
