LiDAR-Based Online 3D Video Object Detection With Graph-Based Message Passing and Spatiotemporal Transformer Attention
Junbo Yin, Jianbing Shen, Chenye Guan, Dingfu Zhou, Ruigang Yang
摘要
Existing LiDAR-based 3D object detectors usually focus on the single-frame detection, while ignoring the spatiotemporal information in consecutive point cloud frames. In this paper, we propose an end-to-end online 3D video object detector that operates on point cloud sequences. The proposed model comprises a spatial feature encoding component and a spatiotemporal feature aggregation component. In the former component, a novel Pillar Message Passing Network (PMPNet) is proposed to encode each discrete point cloud frame. It adaptively collects information for a pillar node from its neighbors by iterative message passing, which effectively enlarges the receptive field of the pillar feature. In the latter component, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU) to aggregate the spatiotemporal information, which enhances the conventional ConvGRU with an attentive memory gating mechanism. AST-GRU contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module, which can emphasize the foreground objects and align the dynamic objects, respectively. Experimental results demonstrate that the proposed 3D video object detector achieves state-of-the-art performance on the large-scale nuScenes benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 被引用 602 次
- Multimodal Virtual Point 3D DetectionTianwei Yin, Xingyi Zhou, Philipp KrähenbühlNeurIPS 2021 · 被引用 379 次
- MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly DetectionYingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton W. T. Fok 等AAAI 2023 · 被引用 221 次
- AFDetV2: Rethinking the Necessity of the Second Stage for Object Detection from Point CloudsYihan Hu, Zhuangzhuang Ding, Runzhou Ge, Wenxin Shao 等AAAI 2022 · 被引用 163 次
- Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point CloudsChaoda Zheng, Xu Yan, Haiming Zhang, Baoyuan Wang 等CVPR 2022 · 被引用 100 次
它引用的顶会 Paper11
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen 等ICCV 2019 · 被引用 840 次
- An Empirical Study of Spatial Attention Mechanisms in Deep NetworksXizhou Zhu, Dazhi Cheng, Zheng Zhang, Stephen Lin 等ICCV 2019 · 被引用 522 次
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera 等ICCV 2019 · 被引用 504 次
- Fast Point R-CNNYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaICCV 2019 · 被引用 440 次
相关 Paper
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim 等AAAI 2023 · 被引用 34 次
- PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionKuan-Chih Huang, Weijie Lyu, Ming-Hsuan Yang, Yi-Hsuan TsaiCVPR 2024
- PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point CloudsJinyu Li, Chenxu Luo, Xiaodong YangCVPR 2023
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu 等ICCV 2025 · 被引用 2 次
- Cross Modal Transformer: Towards Fast and Robust 3D Object DetectionJunjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia 等ICCV 2023 · 被引用 143 次
