SAIL-VOS 3D: A Synthetic Dataset and Baselines for Object Detection and 3D Mesh Reconstruction From Video Data
Yuan-Ting Hu, Jiahong Wang, Raymond A. Yeh, Alexander G. Schwing
Abstract
Extracting detailed 3D information of objects from video data is an important goal for holistic scene understanding. While recent methods have shown impressive results when reconstructing meshes of objects from a single image, results often remain ambiguous as part of the object is unobserved. Moreover, existing image-based datasets for mesh reconstruction don’t permit to study models which integrate temporal information. To alleviate both concerns we present SAIL-VOS 3D: a synthetic video dataset with frame-by-frame mesh annotations which extends SAIL-VOS. We also develop first baselines for reconstruction of 3D meshes from video data via temporal models. We demonstrate efficacy of the proposed baseline on SAIL-VOS 3D and Pix3D, showing that temporal information improves reconstruction quality. Resources and additional information are available at http://sailvos.web.illinois.edu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- CASA: Category-agnostic Skeletal Animal ReconstructionYuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren et al.NeurIPS 2022 · 47 citations
- LRM-Zero: Training Large Reconstruction Models with Synthesized DataDesai Xie, Sai Bi, Zhixin Shu, Kai Zhang et al.NeurIPS 2024 · 36 citations
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- 4D Panoptic Scene Graph GenerationJingkang Yang, Jun Cen, Wenxuan Peng, Shuai Liu et al.NeurIPS 2023 · 33 citations
- Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point CloudsZhao Jin, Yinjie Lei, Naveed Akhtar, Haifeng Li et al.CVPR 2022 · 26 citations
Builds on12
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- PolyGen: An Autoregressive Generative Model of 3D MeshesCharlie Nash, Yaroslav Ganin, S. M. Ali Eslami, Peter W. BattagliaICML 2020 · 339 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 314 citations
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 279 citations
Related papers
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim et al.NeurIPS 2020 · 62 citations
- Deep 3D Mask Volume for View Synthesis of Dynamic ScenesKai-En Lin, Lei Xiao, Feng Liu, Guowei Yang et al.ICCV 2021 · 42 citations
- Unsupervised object-centric video generation and decomposition in 3DPaul Henderson, Christoph H. LampertNeurIPS 2020 · 41 citations
- Clip Fusion with Bi-level Optimization for Human Mesh Reconstruction from Monocular VideosPeng Wu, Xiankai Lu, Jianbing Shen, Yilong YinACM MM 2023 · 18 citations
- WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under OcclusionKhiem Vuong, N. Dinesh Reddy, Robert Tamburo, Srinivasa G. NarasimhanCVPR 2024 · 1 citation
