Optimal Transport-based Labor-free Text Prompt Modeling for Sketch Re-identification
Rui Li, Tingting Ren, Jie Wen, Jinxing Li
摘要
Sketch Re-identification (Sketch Re-ID), which aims to retrieve target person from an image gallery based on a sketch query, is crucial for criminal investigation, law enforcement, and missing person searches. Existing methods aim to alleviate the modality gap by employing semantic metrics constraints or auxiliary modal guidance. However, they incur expensive labor costs and inevitably omit fine-grained modality-consistent information due to the abstraction of sketches. To address this issue, this paper proposes a novel Optimal Transport-based Labor-free Text Prompt Modeling (OLTM) network, which hierarchically extracts coarse-and fine-grained similarity representations guided by textual semantic information without any additional annotations. Specifically, multiple target attributes are flexibly obtained by a pre-trained visual question answering (VQA) model. Subsequently, a text prompt reasoning module employs learnable prompt strategy and optimal transport algorithm to extract discriminative global and local text representations, which serve as a bridge for hierarchical and multi-granularity modal alignment between sketch and image modalities. Additionally, instead of measuring the similarity of two samples by only computing their distance, a novel triplet assignment loss is further proposed, in which the whole data distribution also contributes to optimizing the inter/intra-class distances. Extensive experiments conducted on two public benchmarks consistently demonstrate the robustness and superiority of our OLTM over state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationLinhan Zhou, Shuang Li, Neng Dong, Yonghang Tai 等AAAI 2026 · 被引用 4 次
- Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identificationYajing Zhai, Yawen Zeng, Zhiyong Huang, Zheng Qin 等AAAI 2024 · 被引用 40 次
- Towards Modality-Agnostic Person Re-identification with Descriptive QueryCuiqun Chen, Mang Ye, Ding JiangCVPR 2023
- ASMR: Learning Attribute-Based Person Search with Adaptive Semantic Margin RegularizerBoseung Jeong, Jicheol Park, Suha KwakICCV 2021 · 被引用 29 次
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and AlignmentTong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang 等EMNLP 2025 · 被引用 1 次
