Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng, Kaicheng Yang, Qichuan Ding
摘要
Although Contrastive Language-Image Pretraining (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated vision-language data focused on person-centric images, and (ii) the inherent limitations of global contrastive learning, which struggles to maintain discriminative local features crucial for fine-grained matching while remaining vulnerable to noisy text tokens. This work advances CLIP for person representation learning through synergistic improvements in data curation and model architecture. First, we develop a noise-resistant data construction pipeline that leverages the in-context learning capabilities of MLLMs to automatically filter and caption web-sourced images. This yields WebPerson, a large-scale dataset of 5M high-quality person-centric image-text pairs. Second, we introduce the GA-DMS (Gradient-Attention Guided Dual-Masking Synergetic) framework, which improves cross-modal alignment by adaptively masking noisy textual tokens based on the gradient-attention similarity score. Additionally, we incorporate masked token prediction objectives that compel the model to predict informative text tokens, enhancing fine-grained semantic representation learning. Extensive experiments show that GA-DMS achieves state-of-the-art performance across multiple benchmarks. The data and pre-trained models are released at https://github.com/ Multimodal-Representation-Learning-MRL/ GA-DMS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding LearningTiancheng Gu, Kaicheng Yang, Kaichen Zhang, Xiang An 等AAAI 2026 · 被引用 24 次
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou 等CVPR 2026
它引用的顶会 Paper26
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 被引用 355 次
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin 等ACM MM 2022 · 被引用 197 次
相关 Paper
- PLIP: Language-Image Pre-training for Person Representation LearningJialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu 等NeurIPS 2024 · 被引用 96 次
- RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation ParadigmTiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie 等ACM MM 2025
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng 等EMNLP 2024 · 被引用 11 次
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 等ICML 2025
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong 等ICML 2026 · 被引用 4 次
