Localized Semantic Feature Mixers for Efficient Pedestrian Detection in Autonomous Driving
Abdul Hannan Khan, Mohammed Shariq Nawaz, Andreas Dengel
Abstract
Autonomous driving systems rely heavily on the underlying perception module which needs to be both performant and efficient to allow precise decisions in real-time. Avoiding collisions with pedestrians is of topmost priority in any autonomous driving system. Therefore, pedestrian detection is one of the core parts of such systems' perception modules. Current state-of-the-art pedestrian detectors have two major issues. Firstly, they have long inference times which affect the efficiency of the whole perception module, and secondly, their performance in the case of small and heavily occluded pedestrians is poor. We propose Localized Semantic Feature Mixers (LSFM), a novel, anchor-free pedestrian detection architecture. It uses our novel Super Pixel Pyramid Pooling module instead of the, computationally costly, Feature Pyramid Networks for feature encoding. Moreover, our MLPMixer-based Dense Focal Detection Network is used as a light detection head, reducing computational effort and inference time compared to existing approaches. To boost the performance of the proposed architecture, we adapt and use mixup augmentation which improves the performance, especially in small and heavily occluded cases. We benchmark LSFM against the state-of-the-art on well-established traffic scene pedestrian datasets. The proposed LSFM achieves state-of-the-art performance in Caltech, City Persons, Euro City Persons, and TJU-Traffic-Pedestrian datasets while reducing the inference time on average by 55%. Further, LSFM beats the human baseline for the first time in the history of pedestrian detection. Finally, we conducted a cross-dataset evaluation which proved that our proposed LSFM generalizes well to unseen data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 250cc2a6-fbda-4085-ad23-97aecc6dbd39Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
Related papers
- VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-SupervisionMengyin Liu, Jie Jiang, Chao Zhu, Xu-Cheng YinCVPR 2023
- Learning Hierarchical Graph for Occluded Pedestrian DetectionGang Li, Jian Li, Shanshan Zhang, Jian YangACM MM 2020 · 11 citations
- Variational Pedestrian DetectionYuang Zhang, Huanyu He, Jianguo Li, Yuxi Li et al.CVPR 2021
- Self-Mimic Learning for Small-scale Pedestrian DetectionJialian Wu, Chunluan Zhou, Qian Zhang, Ming Yang et al.ACM MM 2020 · 67 citations
- PedHunter: Occlusion Robust Pedestrian Detector in Crowded ScenesCheng Chi, Shifeng Zhang, Junliang Xing, Zhen Lei et al.AAAI 2020 · 118 citations
