FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search
Xiaoxuan Mu, Haoyu Tang, Han Jiang, Tianyuan Liang, Qinghai Zheng, Jihua Zhu
Abstract
Text-Based Person Search, which aims to retrieve target pedestrian images using natural language descriptions, has garnered significant attention in multimedia research due to its potential in suspect retrieval and missing person identification. While supervised and weakly supervised methods rely on costly annotated training data, unsupervised TBPS eliminates the need for textual descriptions or identity annotations, presenting a more practical paradigm. Current unsupervised TBPS approaches face two primary challenges: 1) Predefined attribute templates for caption generation limit linguistic diversity and real-world adaptability, and 2) Threshold-based sample selection using pre-trained vision-language models (VLMs) introduces noisy pairs due to inadequate pedestrian-specific representation. To address these limitations, we propose FACE, a unified framework featuring Dual-template Caption Generation (DCG) and Adaptive Curriculum Training (ACT). The DCG module generates high-quality captions through complementary flexible-style (natural language) and fixed-style (attribute-enumerated) templates, enhanced by LLM-based noise filtering. The ACT framework progressively refines training through a self-improving loop: initial high-confidence sample selection using VLMs bootstraps the model, while evolving feature representations enable dynamic incorporation of harder samples through curriculum learning. This dual strategy achieves mutual reinforcement between caption quality and model discriminability. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets under unsupervised settings demonstrate that our framework achieves the state-of-the-art performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f97720d0-8a17-46a8-941b-1596bd14b54fRelated papers
- Text-based Person Search without Parallel Image-Text DataYang Bai, Jingyao Wang, Min Cao, Chen Chen et al.ACM MM 2023 · 27 citations
- Unsupervised Cross-Modal Person Search via Progressive Diverse Text GenerationFeng Chen, Jielong He, Yang Liu, Heng Liu et al.ACM MM 2025 · 1 citation
- Test-Time Adaptation for Text-Based Person SearchKai Niu, Liucun Shi, Ke Han, Qinzi Zhao et al.ACM MM 2025
- GPT-ReID: Learning Fine-grained Representation with GPT for Text-based Person RetrievalXudong Wang, Lei Tan, Pingyang Dai, Liujuan Cao et al.ACM MM 2025
- Diverse Person: Customize Your Own Dataset for Text-Based Person SearchZifan Song, Guosheng Hu, Cairong ZhaoAAAI 2024
