Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
Yiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang, Weilin Zhuang, Rongrong Ji
摘要
Text-based person retrieval (TPR) is a challenging task that involves retrieving a specific individual based on a textual description. Despite considerable efforts to bridge the gap between vision and language, the significant differences between these modalities continue to pose a challenge. Previous methods have attempted to align text and image samples in a modal-shared space, but they face uncertainties in optimization directions due to the movable features of both modalities and the failure to account for one-to-many relationships of image-text pairs in TPR datasets. To address this issue, we propose an effective bi-directional one-to-many embedding paradigm that offers a clear optimization direction for each sample, thus mitigating the optimization problem. Additionally, this embedding scheme generates multiple features for each sample without introducing trainable parameters, making it easier to align with several positive samples. Based on this paradigm, we propose a novel Bi-directional one-to-many Embedding Alignment (Beat) model to address the TPR task. Our experimental results demonstrate that the proposed Beat model achieves state-of-the-art performance on three popular TPR datasets, including CUHK-PEDES (65.61 R@1), ICFG-PEDES (58.25 R@1), and RSTPReID (48.10 R@1). Furthermore, additional experiments on MS-COCO, CUB, and Flowers datasets further demonstrate the potential of Beat to be applied to other image-text retrieval tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng 等CVPR 2024 · 被引用 83 次
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang 等ACM MM 2024 · 被引用 16 次
- Chat-Driven Text Generation and Interaction for Person RetrievalZequn Xie, Chuxin Wang, Yeqiang Wang, Sihang Cai 等EMNLP 2025 · 被引用 12 次
- Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person RetrievalZongyi Li, Jianbo Li, Yuxuan Shi, Jiazhong Chen 等AAAI 2025 · 被引用 5 次
- Quota-Calibrated Fine-Grained Alignment with Context-Aware Marginals for Text-based Person RetrievalDongsheng Li, Xinyuan Guo, Huijie Zhang, Pingting Hao 等CVPR 2026
它引用的顶会 Paper24
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
- Scaling Up Vision-Language Pretraining for Image CaptioningXiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang 等CVPR 2022 · 被引用 203 次
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin 等ACM MM 2022 · 被引用 197 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang 等AAAI 2020 · 被引用 182 次
相关 Paper
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang 等AAAI 2024 · 被引用 8 次
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
- Look Before You Leap: Improving Text-based Person Retrieval by Learning A Consistent Cross-modal Common ManifoldZijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan 等ACM MM 2022 · 被引用 125 次
- Test-Time Adaptation for Text-Based Person SearchKai Niu, Liucun Shi, Ke Han, Qinzi Zhao 等ACM MM 2025
- CAIBC: Capturing All-round Information Beyond Color for Text-based Person RetrievalZijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan 等ACM MM 2022 · 被引用 126 次
