Towards Modality-Agnostic Person Re-identification with Descriptive Query
Cuiqun Chen, Mang Ye, Ding Jiang
Abstract
Person re-identification (ReID) with descriptive query (text or sketch) provides an important supplement for general image-image paradigms, which is usually studied in a single cross-modality matching manner, e.g., text-to-image or sketch-to-photo. However, without a camera-captured photo query, it is uncertain whether the text or sketch is available or not in practical scenarios. This motivates us to study a new and challenging modality-agnostic person re-identification problem. Towards this goal, we propose a unified person re-identification (UNIReID) architecture that can effectively adapt to cross-modality and multi-modality tasks. Specifically, UNIReID incorporates a simple dualencoder with task-specific modality learning to mine and fuse visual and textual modality information. To deal with the imbalanced training problem of different tasks in UNIReID, we propose a task-aware dynamic training strategy in terms of task difficulty, adaptively adjusting the training focus. Besides, we construct three multi-modal ReID datasets by collecting the corresponding sketches from photos to support this challenging study. The experimental results on three multi-modal ReID datasets show that our UNIReID greatly improves the retrieval accuracy and generalization ability on different tasks and unseen scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person RetrievalYiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang et al.ACM MM 2023 · 34 citations
- Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDWentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang et al.CVPR 2024 · 34 citations
- ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single ModelJialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin et al.NeurIPS 2025 · 11 citations
- A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-IdentificationYunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep et al.AAAI 2026 · 7 citations
- Optimal Transport-based Labor-free Text Prompt Modeling for Sketch Re-identificationRui Li, Tingting Ren, Jie Wen, Jinxing LiNeurIPS 2024 · 3 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang et al.CVPR 2022 · 313 citations
- Channel Augmented Joint Learning for Visible-Infrared RecognitionMang Ye, Weijian Ruan, Bo Du, Mike Zheng ShouICCV 2021 · 310 citations
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
Related papers
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-IdentificationZhen Sun, Lei Tan, Yunhang Shen, Chengmao Cai et al.ICML 2025
- All in One Framework for Multimodal Re-Identification in the WildHe Li, Mang Ye, Ming Zhang, Bo DuCVPR 2024
- Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Jun Chen, Mang YeICCV 2023 · 53 citations
- Augmented Dual-Contrastive Aggregation Learning for Unsupervised Visible-Infrared Person Re-IdentificationBin Yang, Mang Ye, Jun Chen, Zesen WuACM MM 2022 · 113 citations
- Hierarchical Prompt Learning for Image- and Text-Based Person Re-IdentificationLinhan Zhou, Shuang Li, Neng Dong, Yonghang Tai et al.AAAI 2026 · 4 citations
