Person Tube Retrieval via Language Description
Hehe Fan, Yi Yang
摘要
This paper focuses on the problem of person tube (a sequence of bounding boxes which encloses a person in a video) retrieval using a natural language query. Different from images in person re-identification (re-ID) or person search, besides appearance, person tube contains abundant action and information. We exploit a 2D and a 3D residual networks (ResNets) to extract the appearance and action representation, respectively. To transform tubes and descriptions into a shared latent space where data from the two different modalities can be compared directly, we propose a Multi-Scale Structure Preservation (MSSP) approach. MSSP splits a person tube into several element-tubes on average, whose features are extracted by the two ResNets. Any number of consecutive element-tubes forms a sub-tube. MSSP considers the following constraints for sub-tubes and descriptions in the shared space. 1) Bidirectional ranking. Matching sub-tubes (resp. descriptions) should get ranked higher than incorrect ones for each description (resp. sub-tube). 2) External structure preservation. Sub-tubes (resp. descriptions) from different persons should stay away from each other. 3) Internal structure preservation. Sub-tubes (resp. descriptions) from the same person should be close to each other. Experimental results on person tube retrieval via language description and other two related tasks demonstrate the efficacy of MSSP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- T2VLAD: Global-Local Sequence Alignment for Text-Video RetrievalXiaohan Wang, Linchao Zhu, Yi YangCVPR 2021
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
它引用的顶会 Paper1
相关 Paper
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang 等ACM MM 2022 · 被引用 35 次
- Maskable Retentive Network for Video Moment RetrievalJingjing Hu, Dan Guo, Kun Li, Zhan Si 等ACM MM 2024 · 被引用 7 次
- TVPR: Text-to-Video Person Retrieval and a New BenchmarkXu Zhang, Fan Ni, Guannan Dong, Aichun Zhu 等ACM MM 2024 · 被引用 2 次
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 被引用 69 次
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 被引用 75 次
