Cuboids Revisited: Learning Robust 3D Shape Fitting to Single RGB Images
Florian Kluger, Hanno Ackermann, Eric Brachmann, Michael Ying Yang, Bodo Rosenhahn
摘要
Humans perceive and construct the surrounding world as an arrangement of simple parametric models. In particular, man-made environments commonly consist of volumetric primitives such as cuboids or cylinders. Inferring these primitives is an important step to attain high-level, abstract scene descriptions. Previous approaches directly estimate shape parameters from a 2D or 3D input, and are only able to reproduce simple objects, yet unable to accurately parse more complex 3D scenes. In contrast, we propose a robust estimator for primitive fitting, which can meaningfully abstract real-world environments using cuboids. A RANSAC estimator guided by a neural network fits these primitives to 3D features, such as a depth map. We condition the network on previously detected parts of the scene, thus parsing it one-by-one. To obtain 3D features from a single RGB image, we additionally optimise a feature extraction CNN in an end-to-end manner. However, naively minimising pointto-primitive distances leads to large or spurious cuboids occluding parts of the scene behind. We thus propose an occlusion-aware distance metric correctly handling opaque scenes. The proposed algorithm does not require labourintensive labels, such as cuboid annotations, for training. Results on the challenging NYU Depth v2 dataset demonstrate that the proposed algorithm successfully abstracts cluttered real-world 3D scene layouts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Differentiable Blocks World: Qualitative 3D Decomposition by Rendering PrimitivesTom Monnier, Jake Austin, Angjoo Kanazawa, Alexei A. Efros 等NeurIPS 2023 · 被引用 50 次
- Iterative Superquadric Recomposition of 3D Objects from Multiple ViewsStephan Alaniz, Massimiliano Mancini, Zeynep AkataICCV 2023 · 被引用 21 次
- Coin3D: Controllable and Interactive 3D Assets Generation with Proxy-Guided ConditioningWenqi Dong, Bangbang Yang, Lin Ma, Xiao Liu 等SIGGRAPH 2024 · 被引用 18 次
- PARSAC: Accelerating Robust Multi-Model Fitting with Parallel Sample ConsensusFlorian Kluger, Bodo RosenhahnAAAI 2024 · 被引用 11 次
- Convex Decomposition of Indoor ScenesVaibhav Vavilala, David A. ForsythICCV 2023 · 被引用 11 次
它引用的顶会 Paper10
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna 等ICCV 2019 · 被引用 427 次
- Neural-Guided RANSAC: Learning Where to Sample Model HypothesesEric Brachmann, Carsten RotherICCV 2019 · 被引用 282 次
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang 等ICCV 2019 · 被引用 218 次
相关 Paper
- Single Image 3D Object Estimation with Primitive Graph NetworksQian He, Desen Zhou, Bo Wan, Xuming HeACM MM 2021 · 被引用 1 次
- PlaneRAS: Learning Planar Primitives for 3D Plane RecoveryFang Zhang, Wenzhao Zheng, Linqing Zhao, Zelan Zhu 等ICCV 2025
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- Neural Part Priors: Learning to Optimize Part-Based Object Completion in RGB-D ScansAleksei Bokhovkin, Angela DaiCVPR 2023
- Holistic 3D Human and Scene Mesh Estimation From Single View ImagesZhenzhen Weng, Serena YeungCVPR 2021
