MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
Pengxiang Wu, Siheng Chen, Dimitris N. Metaxas
Abstract
The ability to reliably perceive the environmental states,particularly the existence of objects and their motion behavior, is crucial for autonomous driving. In this work, we propose an efficient deep model, called MotionNet, to jointly perform perception and motion prediction from 3D point clouds. MotionNet takes a sequence of LiDAR sweeps as input and outputs a bird's eye view (BEV) map, which encodes the object category and motion information in each grid cell. The backbone of MotionNet is a novel spatiotemporal pyramid network, which extracts deep spatial and temporal features in a hierarchical fashion. To enforce the smoothness of predictions over both space and time, the training of MotionNet is further regularized with novel spatial and temporal consistency losses. Extensive experiments show that the proposed method overall outperforms the state-of-the-arts, including the latest scene-flowand 3D-object-detection-based methods. This indicates the potential value of the proposed method serving as a backup to the bounding-box-based system, and providing complementary information to the motion planner in autonomous driving IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44ef2136-615b-43e0-b289-30c3fea10c38Cited by top-tier papers34
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
- How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionDingkang Yang, Kun Yang, Yuzheng Wang, Jing Liu et al.NeurIPS 2023 · 160 citations
- EMP: edge-assisted multi-vehicle perceptionXumiao Zhang, Anlan Zhang, Jiachen Sun, Xiao Zhu et al.MobiCom 2021 · 137 citations
- Exploring Simple 3D Multi-Object Tracking for Autonomous DrivingChenxu Luo, Xiaodong Yang, Alan L. YuilleICCV 2021 · 122 citations
Builds on5
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- MeteorNet: Deep Learning on Dynamic 3D Point Cloud SequencesXingyu Liu, Mengyuan Yan, Jeannette BohgICCV 2019 · 225 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim et al.AAAI 2023 · 34 citations
- BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional EnhancementYunlong Wang, Hongyu Pan, Jun Zhu, Yu-Huan Wu et al.CVPR 2022 · 22 citations
- Query-based Temporal Fusion with Explicit Motion for 3D Object DetectionJinghua Hou, Zhe Liu, Dingkang Liang, Zhikang Zou et al.NeurIPS 2023 · 28 citations
- AsyncBEV: Cross-modal flow alignment in Asynchronous 3D Object DetectionShiming Wang, Holger Caesar, Liangliang Nan, Julian F. P. KooijICLR 2026 · 2 citations
- Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesSanmin Kim, Youngseok Kim, In-Jae Lee, Dongsuk KumICCV 2023 · 16 citations
