Stacked Homography Transformations for Multi-View Pedestrian Detection
Liangchen Song, Jialian Wu, Ming Yang, Qian Zhang, Yuan Li, Junsong Yuan
Abstract
Multi-view pedestrian detection aims to predict a bird’s eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we propose a novel Stacked HOmography Transformations (SHOT) approach, which is motivated by approximating projections in 3D world coordinates via a stack of homographies. We first construct a stack of transformations for projecting views to the ground plane at different height levels. Then we design a soft selection module so that the network learns to predict the likelihood of the stack of transformations. Moreover, we provide an in-depth theoretical analysis on constructing SHOT and how well SHOT approximates projections in 3D world coordinates. SHOT is empirically verified to be capable of estimating accurate correspondences from individual views to the BEV map, leading to new state-of-the-art performance on standard evaluation benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be5698e6-321c-4eb9-b9e7-e7b835128e1eCited by top-tier papers11
- From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera CalibrationZekun Qian, Ruize Han, Wei Feng, Song WangCVPR 2024 · 8 citations
- Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution WeightingQi Zhang, Yunfei Gong, Daijie Chen, Antoni B. Chan et al.AAAI 2024 · 7 citations
- Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic DatasetSithu Aung, Min-Cheol Sagong, Junghyun ChoAAAI 2025 · 5 citations
- Unsupervised Multi-view Pedestrian DetectionMengyin Liu, Chao Zhu, Shiqi Ren, Xu-Cheng YinACM MM 2024 · 4 citations
- MVTrajecter: Multi-View Pedestrian Tracking With Trajectory Motion Cost and Trajectory Appearance CostTaiga Yamane, Ryo Masumura, Satoshi Suzuki, Shota OrihashiICCV 2025 · 2 citations
Builds on6
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- 3D Crowd Counting via Multi-View Fusion with 3D Gaussian KernelsQi Zhang, Antoni B. ChanAAAI 2020 · 41 citations
- Simultaneous Multi-View Instance Detection With Learned Geometric Soft-ConstraintsAhmed Samy Nassar, Sébastien Lefèvre, Jan Dirk WegnerICCV 2019 · 29 citations
- Epipolar TransformersYihui He, Rui Yan, Katerina Fragkiadaki, Shoou-I YuCVPR 2020
- Cross-View Cross-Scene Multi-View Crowd CountingQi Zhang, Wei Lin, Antoni B. ChanCVPR 2021
Related papers
- PandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose EstimationAbdallah Benzine, Florian Chabot, Bertrand Luvison, Quoc Cuong Pham et al.CVPR 2020
- BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention NetworksXiaowei Chi, Jiaming Liu, Ming Lu, Rongyu Zhang et al.CVPR 2023
- Discriminative Spatial Feature Learning for Person Re-IdentificationPeixi Peng, Yonghong Tian, Yangru Huang, Xiangqian Wang et al.ACM MM 2020 · 5 citations
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionZhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue et al.ICCV 2023 · 63 citations
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
