Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly Search
Shuyu Yang, Yaxiong Wang, Li Zhu, Zhedong Zheng
摘要
Text-based person search aims to retrieve specific individuals across camera networks using natural language descriptions. However, current benchmarks often exhibit biases towards common actions like walking or standing, neglecting the critical need for identifying abnormal behaviors in real-world scenarios. To meet such demands, we propose a new task, text-based person anomaly search, locating pedestrians engaged in both routine or anomalous activities via text. To enable the training and evaluation of this new task, we construct a large-scale image-text Pedestrian Anomaly Behavior (PAB) benchmark, featuring a broad spectrum of actions, e.g., running, performing, playing soccer, and the corresponding anomalies, e.g., lying, being hit, and falling of the same identity. The training set of PAB comprises 1,013,605 synthesized image-text pairs of both normalities and anomalies, while the test set includes 1,978 real-world image-text pairs. To validate the potential of PAB, we introduce a cross-modal pose-aware framework, which integrates human pose patterns with identity-based hard negative pair sampling. Extensive experiments on the proposed benchmark show that synthetic training data facilitates the fine-grained behavior retrieval, and the proposed pose-aware method arrives at 84.93% recall@1 accuracy, surpassing other competitive methods. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/CMP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- From Data Deluge to Data Curation: A Filtering-WoRA Paradigm for Efficient Text-based Person SearchJintao Sun, Hao Fei, Gangyi Ding, Zhedong ZhengWWW 2025 · 被引用 28 次
- UniAD: Integrating Geometric and Semantic Cues for Unified Anomaly DetectionXiaodong Wang, Hongmin Hu, Fei Yan, Junwen Lu 等ACM MM 2025 · 被引用 2 次
- Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person SearchJiahao Zhang, Shaofei Huang, Yaxiong Wang, Zhedong ZhengSIGIR 2026
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsYan Zeng, Xinsong Zhang, Hang LiICML 2022 · 被引用 371 次
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
相关 Paper
- UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityJialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang 等CVPR 2024 · 被引用 45 次
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang 等AAAI 2020 · 被引用 182 次
- Graph Embedded Pose Clustering for Anomaly DetectionAmir Markovitz, Gilad Sharir, Itamar Friedman, Lihi Zelnik-Manor 等CVPR 2020
- PIAD: Pose and Illumination agnostic Anomaly DetectionKaichen Yang, Junjie Cao, Zeyu Bai, Zhixun Su 等CVPR 2025
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye 等AAAI 2024 · 被引用 111 次
