BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional Enhancement
Yunlong Wang, Hongyu Pan, Jun Zhu, Yu-Huan Wu, Xin Zhan, Kun Jiang, Diange Yang
Abstract
Determining the motion behavior of inexhaustible categories of traffic participants is critical for autonomous driving. In recent years, there has been a rising concern in performing class-agnostic motion prediction directly from the captured sensor data, like LiDAR point clouds or the combination of point clouds and images. Current motion prediction frameworks tend to perform joint semantic segmentation and motion prediction and face the trade-off between the performance of these two tasks. In this paper, we propose a novel Spatial-Temporal Integrated network with Bidirectional Enhancement, BE-STI, to improve the temporal motion prediction performance by spatial semantic features, which points out an efficient way to combine semantic segmentation and motion prediction. Specifically, we propose to enhance the spatial features of each individual point cloud with the similarity among temporal neighboring frames and enhance the global temporal features with the spatial difference among non-adjacent frames in a coarse-to-fine fashion. Extensive experiments on nuScenes and Waymo Open Dataset show that our proposed framework outperforms all state-of-the-art LiDAR-based and RGB+LiDAR-based methods with remarkable margins by using only point clouds as input. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> The code will be released at https://github.com/be-sti/be-sti.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d122d63-6e49-4005-a0fc-c2f2fe6018ccCited by top-tier papers7
- Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMixKewei Wang, Yizheng Wu, Zhiyu Pan, Xingyi Li et al.AAAI 2024 · 11 citations
- Self-Supervised Bird's Eye View Motion Prediction with Cross-Modality SignalsShaoheng Fang, Zuhong Liu, Mingyu Wang, Chenxin Xu et al.AAAI 2024 · 8 citations
- Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency RegularizationsKewei Wang, Yizheng Wu, Jun Cen, Zhiyu Pan et al.CVPR 2024 · 3 citations
- PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field PriorsKangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo et al.ICCV 2025
- TBP-Former: Learning Temporal Bird's-Eye-View Pyramid for Joint Perception and Prediction in Vision-Centric Autonomous DrivingShaoheng Fang, Zi Wang, Yiqi Zhong, Junhao Ge et al.CVPR 2023
Builds on12
- Spatial-Temporal Relation Networks for Multi-Object TrackingJiarui Xu, Yue Cao, Zheng Zhang, Han HuICCV 2019 · 260 citations
- Robust Multi-Modality Multi-Object TrackingWenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang et al.ICCV 2019 · 221 citations
- PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object DetectionShaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang et al.CVPR 2020
- PV-RAFT: Point-Voxel Correlation Fields for Scene Flow Estimation of Point CloudsYi Wei, Ziyi Wang, Yongming Rao, Jiwen Lu et al.CVPR 2021
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View MapsPengxiang Wu, Siheng Chen, Dimitris N. MetaxasCVPR 2020
Related papers
- Query-based Temporal Fusion with Explicit Motion for 3D Object DetectionJinghua Hou, Zhe Liu, Dingkang Liang, Zhikang Zou et al.NeurIPS 2023 · 28 citations
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim et al.AAAI 2023 · 34 citations
- MemorySeg: Online LiDAR Semantic Segmentation with a Latent MemoryEnxu Li, Sergio Casas, Raquel UrtasunICCV 2023 · 26 citations
- STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory PredictionZhishuai Zhang, Jiyang Gao, Junhua Mao, Yukai Liu et al.CVPR 2020
- TPNet: Trajectory Proposal Network for Motion PredictionLiangji Fang, Qinhong Jiang, Jianping Shi, Bolei ZhouCVPR 2020
