Generalizable Pedestrian Detection: The Elephant in the Room
Irtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram, Ling Shao
Abstract
Pedestrian detection is used in many vision based applications ranging from video surveillance to autonomous driving. Despite achieving high performance, it is still largely unknown how well existing detectors generalize to unseen data. This is important because a practical detector should be ready to use in various scenarios in applications. To this end, we conduct a comprehensive study in this paper, using a general principle of direct crossdataset evaluation. Through this study, we find that existing state-of-the-art pedestrian detectors, though perform quite well when trained and tested on the same dataset, generalize poorly in cross dataset evaluation. We demonstrate that there are two reasons for this trend. Firstly, their designs (e.g. anchor settings) may be biased towards popular benchmarks in the traditional single-dataset training and test pipeline, but as a result largely limit their generalization capability. Secondly, the training source is generally not dense in pedestrians and diverse in scenarios. Under direct cross-dataset evaluation, surprisingly, we find that a general purpose object detector, without pedestriantailored adaptation in design, generalizes much better compared to existing state-of-the-art pedestrian detectors. Furthermore, we illustrate that diverse and dense datasets, collected by crawling the web, serve to be an efficient source of pre-training for pedestrian detection. Accordingly, we propose a progressive training pipeline and find that it works well for autonomous-driving oriented pedestrian detection. Consequently, the study conducted in this paper suggests that more emphasis should be put on cross-dataset evaluation for the future design of generalizable pedestrian detectors. Code and models can be accessed at https: //github.com/hasanirtiza/Pedestron .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32443637-fb35-4387-a3d0-e25393d18928Cited by top-tier papers16
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
- Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and BeyondYi Yu, Wenhan Yang, Yap-Peng Tan, Alex C. KotCVPR 2022 · 53 citations
- Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier SearchBicheng Guo, Tao Chen, Shibo He, Haoyu Liu et al.ACM MM 2022 · 21 citations
- BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online PoliciesThomas Verelst, Tinne TuytelaarsICCV 2021 · 18 citations
- Unified Visual Relationship Detection with Vision and Language ModelsLong Zhao, Liangzhe Yuan, Boqing Gong, Yin Cui et al.ICCV 2023 · 13 citations
Builds on6
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer et al.ICCV 2019 · 216 citations
- Self-Mimic Learning for Small-scale Pedestrian DetectionJialian Wu, Chunluan Zhou, Qian Zhang, Ming Yang et al.ACM MM 2020 · 67 citations
- Bridging the Gap Between Detection and Tracking: A Unified ApproachLianghua Huang, Xin Zhao, Kaiqi HuangICCV 2019 · 40 citations
- Box Guided Convolution for Pedestrian DetectionJinpeng Li, Shengcai Liao, Hangzhi Jiang, Ling ShaoACM MM 2020 · 28 citations
- Temporal-Context Enhanced Detection of Heavily Occluded PedestriansJialian Wu, Chunluan Zhou, Ming Yang, Qian Zhang et al.CVPR 2020
Related papers
- VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-SupervisionMengyin Liu, Jie Jiang, Chao Zhu, Xu-Cheng YinCVPR 2023
- Unsupervised Adaptation from Repeated Traversals for Autonomous DrivingYurong You, Cheng Perng Phoo, Katie Luo, Travis Zhang et al.NeurIPS 2022 · 16 citations
- HumanBench: Towards General Human-Centric Perception with Projector Assisted PretrainingShixiang Tang, Cheng Chen, Qingsong Xie, Meilin Chen et al.CVPR 2023
- Rethinking Open-World Object Detection in Autonomous Driving ScenariosZeyu Ma, Yang Yang, Guoqing Wang, Xing Xu et al.ACM MM 2022 · 39 citations
- Learning to Detect Mobile Objects from LiDAR Scans Without LabelsYurong You, Katie Luo, Cheng Perng Phoo, Wei-Lun Chao et al.CVPR 2022 · 33 citations
