To The Point: Correspondence-driven monocular 3D category reconstruction
Filippos Kokkinos, Iasonas Kokkinos
Abstract
We present To The Point (TTP), a method for reconstructing 3D objects from a single image using 2D to 3D correspondences learned from weak supervision. We recover a 3D shape from a 2D image by first regressing the 2D positions corresponding to the 3D template vertices and then jointly estimating a rigid camera transform and non-rigid template deformation that optimally explain the 2D positions through the 3D shape projection. By relying on 3D-2D correspondences we use a simple per-sample optimization problem to replace CNN-based regression of camera pose and non-rigid deformation and thereby obtain substantially more accurate 3D reconstructions. We treat this optimization as a differentiable layer and train the whole system in an end-to-end manner. We report systematic quantitative improvements on multiple categories and provide qualitative results comprising diverse shape, pose and texture prediction examples. Project website: https: //fkokkinos.github.io/to_the_point/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84124e8e-bffc-4062-b0d4-27df2aa12292Cited by top-tier papers12
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan et al.CVPR 2022 · 113 citations
- Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D CameraHongrui Cai, Wanquan Feng, Xuetao Feng, Yan Wang et al.NeurIPS 2022 · 83 citations
- PPR: Physically Plausible Reconstruction from Monocular VideosGengshan Yang, Shuo Yang, John Z. Zhang, Zachary Manchester et al.ICCV 2023 · 41 citations
- Do It Yourself: Learning Semantic Correspondence from Pseudo-LabelsOlaf Dünkel, Thomas Wimmer, Christian Theobalt, Christian Rupprecht et al.ICCV 2025 · 4 citations
- SAOR: Single-View Articulated Object ReconstructionMehmet Aygün, Oisin Mac AodhaCVPR 2024
Builds on11
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From MotionDavid Novotný, Nikhila Ravi, Benjamin Graham, Natalia Neverova et al.ICCV 2019 · 126 citations
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 104 citations
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 72 citations
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim et al.NeurIPS 2020 · 62 citations
Related papers
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Topologically-Aware Deformation Fields for Single-View 3D ReconstructionShivam Duggal, Deepak PathakCVPR 2022 · 30 citations
- Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstructionDavid Novotný, Roman Shapovalov, Andrea VedaldiNeurIPS 2020 · 9 citations
- ShapeClipper: Scalable 3D Shape Learning from Single-View Images via Geometric and CLIP-Based ConsistencyZixuan Huang, Varun Jampani, Anh Thai, Yuanzhen Li et al.CVPR 2023
- Unsupervised Template-assisted Point Cloud Shape Correspondence NetworkJiacheng Deng, Jiahao Lu, Tianzhu ZhangCVPR 2024
