Effectiveness of Vision Transformer for Fast and Accurate Single-Stage Pedestrian Detection
Jing Yuan, Panagiotis Barmpoutis, Tania Stathaki
Abstract
Vision transformers have demonstrated remarkable performance on a variety of computer vision tasks. In this paper, we illustrate the effectiveness of the de-formable vision transformer for single-stage pedestrian detection and propose a spatial and multi-scale feature enhancement module, which aims to achieve the optimal balance between speed and accuracy. Performance improvement with vision transformers on various commonly used single-stage structures is demonstrated. The design of the proposed architecture is investigated in depth. Comprehensive comparisons with state-of-the-art single-and two-stage detectors on different pedestrian datasets are performed. The proposed detector achieves leading performance on Caltech and Citypersons datasets among single-and two-stage methods using fewer parameters than the baseline. The log-average miss rates for Reasonable and Heavy are decreased to 2.6% and 28.0% on the Caltech test set, and 10.9% and 38.6% on the Citypersons validation set, respectively. The proposed method outperforms SOTA two-stage detectors in the Heavy subset on the Citypersons validation set with considerably faster inference speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa2cdf82-5006-4817-892b-338e6b2115beBuilds on11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer et al.ICCV 2019 · 216 citations
- PedHunter: Occlusion Robust Pedestrian Detector in Crowded ScenesCheng Chi, Shifeng Zhang, Junliang Xing, Zhen Lei et al.AAAI 2020 · 118 citations
Related papers
- Self-Mimic Learning for Small-scale Pedestrian DetectionJialian Wu, Chunluan Zhou, Qian Zhang, Ming Yang et al.ACM MM 2020 · 67 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- ViDT: An Efficient and Effective Fully Transformer-based Object DetectorHwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani et al.ICLR 2022 · 96 citations
- Group Vision TransformerYaopeng Peng, Milan Sonka, Danny Z. ChenACM MM 2024
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
