PanoRecon: Real-Time Panoptic 3D Reconstruction from Monocular Video
Dong Wu, Zike Yan, Hongbin Zha
摘要
We introduce the Panoptic 3D Reconstruction task, a unified and holistic scene understanding task for a monocular video. And we present PanoRecon - a novel framework to address this new task, which realizes an online geometry reconstruction alone with dense semantic and instance labeling. Specifically, PanoRecon incrementally performs panoptic 3D reconstruction for each video fragment consisting of multiple consecutive key frames, from a volumetric feature representation using feed-forward neural networks. We adopt a depth-guided back-projection strategy to sparse and purify the volumetric feature representation. We further introduce a voxel clustering module to get object instances in each local fragment, and then design a tracking and fusion algorithm for the integration of instances from different fragments to ensure temporal co-herence. Such design enables our PanoRecon to yield a coherent and accurate panoptic 3D reconstruction. Exper-iments on ScanNetV2 demonstrate a very competitive geometry reconstruction result compared with state-of-the-art reconstruction methods, as well as promising 3D panoptic segmentation result with only RGB input, while being real-time. Code is available at: https://github.com/Riser6/PanoRecon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- PanSt3R: Multi-View Consistent Panoptic SegmentationLojze Zust, Yohann Cabon, Juliette Marrie, Leonid Antsfeld 等ICCV 2025 · 被引用 5 次
- OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian SplattingHongjia Zhai, Qi Zhang, Xiaokun Pan, Xiyu Zhang 等CVPR 2026 · 被引用 3 次
- LIRA: Reasoning Reconstruction via Multimodal Large Language ModelsZhen Zhou, Tong Wang, Yunkai Ma, Xiao Tan 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper24
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 被引用 551 次
- Learning Object-Compositional Neural Radiance Field for Editable Scene RenderingBangbang Yang, Yinda Zhang, Yinghao Xu, Yijin Li 等ICCV 2021 · 被引用 305 次
- Hierarchical Aggregation for 3D Instance SegmentationShaoyu Chen, Jiemin Fang, Qian Zhang, Wenyu Liu 等ICCV 2021 · 被引用 211 次
相关 Paper
- PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosYiming Xie, Matheus Gadelha, Fengting Yang, Xiaowei Zhou 等CVPR 2022 · 被引用 32 次
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 被引用 106 次
- NeuralRecon: Real-Time Coherent 3D Reconstruction From Monocular VideoJiaming Sun, Yiming Xie, Linghao Chen, Xiaowei Zhou 等CVPR 2021
- MGNet: Monocular Geometric Scene Understanding for Autonomous DrivingMarkus Schön, Michael Buchholz, Klaus DietmayerICCV 2021 · 被引用 60 次
- DG-Recon: Depth-Guided Neural 3D Scene ReconstructionJihong Ju, Ching Wei Tseng, Oleksandr Bailo, Georgi Dikov 等ICCV 2023 · 被引用 21 次
