UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View
Zequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen, Xi Li
Abstract
Bird’s eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusion and merges them into a unified mathematical formulation. The unified fusion could not only provide a new perspective on BEV fusion but also brings new capabilities. With the proposed unified spatial-temporal fusion, our method could support long-range fusion, which is hard to achieve in conventional BEV methods. Moreover, the BEV fusion in our work is temporal-adaptive and the weights of temporal fusion are learnable. In contrast, conventional methods mainly use fixed and equal weights for temporal fusion. Besides, the proposed unified fusion could avoid information lost in conventional BEV fusion methods and make full use of features. Extensive experiments and ablation studies on the NuScenes dataset show the effectiveness of the proposed method and our method gains the state-of-the-art performance in the map and vehicle segmentation task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 103625f3-e461-437f-980d-ca80bb479dceCited by top-tier papers10
- MGMap: Mask-Guided Learning for Online Vectorized HD Map ConstructionXiaolu Liu, Song Wang, Wentong Li, Ruizi Yang et al.CVPR 2024 · 36 citations
- BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionWenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan et al.CVPR 2024 · 17 citations
- Extend Your Own Correspondences: Unsupervised Distant Point Cloud Registration by Progressive Distance ExtensionQuan Liu, Hongzi Zhu, Zhenxi Wang, Yunsong Zhou et al.CVPR 2024 · 15 citations
- Failure Modes for Deep Learning-Based Online Mapping: How to Measure and Address ThemMichael Hubbertz, Qi Han, Tobias MeisenCVPR 2026 · 1 citation
- Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation NetworkKunpeng Wang, Keke Chen, Chenglong Li, Zhengzheng Tu et al.AAAI 2025 · 1 citation
Builds on4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object DetectionJisong Kim, Minjae Seong, Jun Won ChoiNeurIPS 2024 · 27 citations
- PointBeV: A Sparse Approach to BeV PredictionsLoïck Chambon, Éloi Zablocki, Mickaël Chen, Florent Bartoccioni et al.CVPR 2024
- MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element ExpertDapeng Zhang, Dayu Chen, Peng Zhi, Yinda Chen et al.AAAI 2025 · 3 citations
- GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion PerceptionXiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang et al.ICLR 2026 · 1 citation
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
