Learning 3D Object Shape and Layout without 3D Supervision
Georgia Gkioxari, Nikhila Ravi, Justin Johnson
摘要
A 3D scene consists of a set of objects, each with a shape and a layout giving their position in space. Understanding 3D scenes from 2D images is an important goal, with ap-plications in robotics and graphics. While there have been recent advances in predicting 3D shape and layout from a single image, most approaches rely on 3D ground truth for training which is expensive to collect at scale. We overcome these limitations and propose a method that learns to predict 3D shape and layout for objects without any ground truth shape or layout information: instead we rely on multi-view images with 2D supervision which can more easily be col-lected at scale. Through extensive experiments on ShapeNet, Hypersim, and ScanNet we demonstrate that our approach scales to large datasets of realistic images, and compares favorably to methods relying on 3D ground truth. On Hy-persim and ScanNet where reliable 3D ground truth is not available, our approach outperforms supervised approaches trained on smaller and less diverse datasets. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Project page https://gkioxari.github.io/usl/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- CAST: Component-Aligned 3D Scene Reconstruction from an RGB ImageKaixin Yao, Longwen Zhang, Xinhao Yan, Yan Zeng 等SIGGRAPH 2025 · 被引用 30 次
- Uni-3D: A Universal Model for Panoptic 3D Scene ReconstructionXiang Zhang, Zeyuan Chen, Fangyin Wei, Zhuowen TuICCV 2023 · 被引用 24 次
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui 等CVPR 2026 · 被引用 10 次
- DepR: Depth Guided Single-View Scene Reconstruction with Instance-Level DiffusionQingcheng Zhao, Xiang Zhang, Haiyang Xu, Zeyuan Chen 等ICCV 2025 · 被引用 3 次
- Monocular Human-Object Reconstruction in the WildChaofan Huo, Ye Shi, Jingya WangACM MM 2024 · 被引用 2 次
它引用的顶会 Paper8
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From MotionDavid Novotný, Nikhila Ravi, Benjamin Graham, Natalia Neverova 等ICCV 2019 · 被引用 126 次
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 被引用 104 次
- Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes From a Single ImageYinyu Nie, Xiaoguang Han, Shihui Guo, Yujian Zheng 等CVPR 2020
相关 Paper
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- ShapeClipper: Scalable 3D Shape Learning from Single-View Images via Geometric and CLIP-Based ConsistencyZixuan Huang, Varun Jampani, Anh Thai, Yuanzhen Li 等CVPR 2023
- 3D Scene Painting via Semantic Image SynthesisJaebong Jeong, Janghun Jo, Sunghyun Cho, Jaesik ParkCVPR 2022 · 被引用 3 次
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva 等ICCV 2025 · 被引用 4 次
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung 等CVPR 2020
