CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion
Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo, Ping Li, Lingyu Liang, Kairui Yang, Di Lin
摘要
Semantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionLubo Wang, Di Lin, Kairui Yang, Ruonan Liu 等NeurIPS 2024 · 被引用 14 次
- Unleashing Network Potentials for Semantic Scene CompletionFengyun Wang, Qianru Sun, Dong Zhang, Jinhui TangCVPR 2024 · 被引用 3 次
- Multi-modal Frequency Decomposition Network for Semantic Scene CompletionDie Zuo, Lubo Wang, Ruonan Liu, Qing Guo 等CVPR 2026
- RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving ScenesYipeng Wu, Xin Wang, Chenghan Yang, Chong Wang 等CVPR 2026
- Point-based Instance Completion with Scene ConstraintsWesley Khademi, Fuxin LiICLR 2025
它引用的顶会 Paper11
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 被引用 251 次
- Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical VoxelizationQi Chen, Lin Sun, Ernest Cheung, Alan L. YuilleNeurIPS 2020 · 被引用 124 次
- VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial AttentionShengheng Deng, Zhihao Liang, Lin Sun, Kui JiaCVPR 2022 · 被引用 92 次
- Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionPingping Zhang, Wei Liu, Yinjie Lei, Huchuan Lu 等ICCV 2019 · 被引用 79 次
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao 等AAAI 2020 · 被引用 68 次
相关 Paper
- Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionZhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu 等NeurIPS 2024 · 被引用 73 次
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao 等CVPR 2023
- H2GFormer: Horizontal-to-Global Voxel Transformer for 3D Semantic Scene CompletionYu Wang, Chao TongAAAI 2024 · 被引用 33 次
- Point Cloud Semantic Scene Completion with Prototype-Guided TransformerChenghao Fang, Jianqing Liang, Jiye Liang, Zijin Du 等AAAI 2026
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 被引用 354 次
