Fine-Grained Image Retrieval via Dual-Vision Adaptation
Xin Jiang, Meiqi Cao, Hao Tang, Fei Shen, Zechao Li
摘要
Fine-Grained Image Retrieval (FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adapted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA performs well on three fine-grained datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum KnowledgeXin Jiang, Hao Tang, Meiqi Cao, Junyao Gao 等CVPR 2026
- Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion PredictionZheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao 等AAAI 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- BlockMix: Meta Regularization and Self-Calibrated Inference for Metric-Based Meta-LearningHao Tang, Zechao Li, Zhimao Peng, Jinhui TangACM MM 2020 · 被引用 120 次
相关 Paper
- Fine-Grained Retrieval Prompt TuningShijie Wang, Jianlong Chang, Zhihui Wang, Haojie Li 等AAAI 2023 · 被引用 27 次
- Adversarial Reconstruction Feedback for Robust Fine-Grained GeneralizationShijie Wang, Jian Shi, Haojie LiICCV 2025 · 被引用 2 次
- DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesXin Jiang, Hao Tang, Rui Yan, Jinhui Tang 等ACM MM 2024 · 被引用 18 次
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang 等CVPR 2026 · 被引用 2 次
- LLM-Assisted Entropy-Based Adaptive Distillation for Unsupervised Fine-Grained Visual Representation LearningJianfeng Dong, Danfeng Luo, Daizong Liu, Jie Sun 等ICCV 2025 · 被引用 1 次
