UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View
Zequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen, Xi Li
摘要
Bird’s eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusion and merges them into a unified mathematical formulation. The unified fusion could not only provide a new perspective on BEV fusion but also brings new capabilities. With the proposed unified spatial-temporal fusion, our method could support long-range fusion, which is hard to achieve in conventional BEV methods. Moreover, the BEV fusion in our work is temporal-adaptive and the weights of temporal fusion are learnable. In contrast, conventional methods mainly use fixed and equal weights for temporal fusion. Besides, the proposed unified fusion could avoid information lost in conventional BEV fusion methods and make full use of features. Extensive experiments and ablation studies on the NuScenes dataset show the effectiveness of the proposed method and our method gains the state-of-the-art performance in the map and vehicle segmentation task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- MGMap: Mask-Guided Learning for Online Vectorized HD Map ConstructionXiaolu Liu, Song Wang, Wentong Li, Ruizi Yang 等CVPR 2024 · 被引用 36 次
- BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionWenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan 等CVPR 2024 · 被引用 17 次
- Extend Your Own Correspondences: Unsupervised Distant Point Cloud Registration by Progressive Distance ExtensionQuan Liu, Hongzi Zhu, Zhenxi Wang, Yunsong Zhou 等CVPR 2024 · 被引用 15 次
- Failure Modes for Deep Learning-Based Online Mapping: How to Measure and Address ThemMichael Hubbertz, Qi Han, Tobias MeisenCVPR 2026 · 被引用 1 次
- Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation NetworkKunpeng Wang, Keke Chen, Chenglong Li, Zhengzheng Tu 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li 等ICCV 2021 · 被引用 404 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
相关 Paper
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object DetectionJisong Kim, Minjae Seong, Jun Won ChoiNeurIPS 2024 · 被引用 27 次
- PointBeV: A Sparse Approach to BeV PredictionsLoïck Chambon, Éloi Zablocki, Mickaël Chen, Florent Bartoccioni 等CVPR 2024
- MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element ExpertDapeng Zhang, Dayu Chen, Peng Zhi, Yinda Chen 等AAAI 2025 · 被引用 3 次
- GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion PerceptionXiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang 等ICLR 2026 · 被引用 1 次
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo 等ACM MM 2023 · 被引用 5 次
