TBP-Former: Learning Temporal Bird's-Eye-View Pyramid for Joint Perception and Prediction in Vision-Centric Autonomous Driving
Shaoheng Fang, Zi Wang, Yiqi Zhong, Junhao Ge, Siheng Chen
Abstract
Vision-centric joint perception and prediction (PnP) has become an emerging trend in autonomous driving research. It predicts the future states of the traffic participants in the surrounding environment from raw RGB images. However, it is still a critical challenge to synchronize features obtained at multiple camera views and timestamps due to inevitable geometric distortions and further exploit those spatial-temporal features. To address this issue, we propose a temporal bird's-eye-view pyramid transformer (TBP-Former) for vision-centric PnP, which includes two novel designs. First, a pose-synchronized BEV encoder is proposed to map raw image inputs with any camera pose at any time to a shared and synchronized BEV space for better spatial-temporal synchronization. Second, a spatialtemporal pyramid transformer is introduced to comprehensively extract multi-scale BEV features and predict future BEV states with the support of spatial priors. Extensive experiments on nuScenes dataset show that our proposed framework overall outperforms all state-of-theart vision-based prediction methods. Code is available at: https://github.com/MediaBrain-SJTU/TBP-Former
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d7f3f23-2c15-44f3-a98d-2cb1037f4367Cited by top-tier papers13
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
- Joint-Relation Transformer for Multi-Person Motion PredictionQingyao Xu, Weibo Mao, Jingze Gong, Chenxin Xu et al.ICCV 2023 · 24 citations
- CaDeT: A Causal Disentanglement Approach for Robust Trajectory Prediction in Autonomous DrivingMozhgan Pourkeshavarz, Junrui Zhang, Amir RasouliCVPR 2024 · 15 citations
- Adversarial Backdoor Attack by Naturalistic Data Poisoning on Trajectory Prediction in Autonomous DrivingMozhgan Pourkeshavarz, Mohammad Sabokrou, Amir RasouliCVPR 2024 · 14 citations
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence MapsYue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong et al.NeurIPS 2022 · 537 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
Related papers
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy PredictionYuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou et al.CVPR 2023
- BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic SegmentationCong Pan, Yonghao He, Junran Peng, Qian Zhang et al.CVPR 2023
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object PredictionZhuofan Zong, Dongzhi Jiang, Guanglu Song, Zeyue Xue et al.ICCV 2023 · 63 citations
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
