SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image
Dimitrije Antic, Georgios Paschalidis, Shashank Tripathi, Theo Gevers, Sai Kumar Dwivedi, Dimitrios Tzionas
摘要
Recovering 3D object pose and shape from a single image is a challenging and ill-posed problem. This is due to strong (self-)occlusions, depth ambiguities, the vast intra- and inter-class shape variance, and the lack of 3D ground truth for natural images. Existing deep-network methods are trained on synthetic datasets to predict 3D shapes, so they often struggle generalizing to real-world images. Moreover, they lack an explicit feedback loop for refining noisy estimates, and primarily focus on geometry without directly considering pixel alignment. To tackle these limitations, we develop a novel render-and-compare optimization framework, called SDFit. This has three key innovations: First, it uses a learned category-specific and morphable signed-distance-function (mSDF) model, and fits this to an image by iteratively refining both 3D pose and shape. The mSDF robustifies inference by constraining the search on the manifold of valid shapes, while allowing for arbitrary shape topologies. Second, SDFit retrieves an initial 3D shape that likely matches the image, by exploiting foundational models for efficient look-up into 3D shape databases. Third, SDFit initializes pose by establishing rich 2D-3D correspondences between the image and the mSDF through foundational features. We evaluate SDFit on three image datasets, i.e., Pix3D, Pascal3D+, and COMIC. SDFit performs on par with SotA feed-forward networks for unoccluded images and common poses, but is uniquely robust to occlusions and uncommon poses. Moreover, it requires no retraining for unseen images. Thus, SDFit contributes new insights for generalizing in the wild. Code is available at https://anticdimi.github.io/sdfit.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM KnowledgeYumeng He, Ying Jiang, Jiayin Lu, Yin Yang 等CVPR 2026 · 被引用 6 次
- TriDi: Trilateral Diffusion of 3D Humans, Objects, and InteractionsIlia A. Petrov, Riccardo Marin, Julian Chibane, Gerard Pons-MollICCV 2025 · 被引用 1 次
- PICO: Reconstructing 3D People In Contact with ObjectsAlpár Cseke, Shashank Tripathi, Sai Kumar Dwivedi, Arjun S. Lakshmipathy 等CVPR 2025
- RHINO: Reconstructing Human Interactions with Novel Objects from Monocular VideosLixin Xue, Chengwei Zheng, Georgios Paschalidis, Chen Guo 等CVPR 2026
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri 等CVPR 2025
相关 Paper
- SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape OptimizationYue Jiang, Dantong Ji, Zhizhong Han, Matthias ZwickerCVPR 2020
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui 等CVPR 2026 · 被引用 10 次
- GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the WildAlexander Grabner, Peter M. Roth, Vincent LepetitICCV 2019 · 被引用 21 次
- GenSDF: Two-Stage Learning of Generalizable Signed Distance FunctionsGene Chou, Ilya Chugunov, Felix HeideNeurIPS 2022 · 被引用 46 次
- SeSDF: Self-Evolved Signed Distance Field for Implicit 3D Clothed Human ReconstructionYukang Cao, Kai Han, Kwan-Yee K. WongCVPR 2023
