CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects
Nick Heppert, Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Rares Andrei Ambrus, Jeannette Bohg, Abhinav Valada, Thomas Kollar
Abstract
We present CARTO, a novel approach for reconstructing multiple articulated objects from a single stereo RGB observation. We use implicit object-centric representations and learn a single geometry and articulation decoder for multiple object categories. Despite training on multiple categories, our decoder achieves a comparable reconstruction accuracy to methods that train bespoke decoders separately for each category. Combined with our stereo image encoder we infer the 3D shape, 6D pose, size, joint type, and the joint state of multiple unknown objects in a single forward pass. Our method achieves a 20.4% absolute improvement in mAP 3D IOU50 for novel instances when compared to a two-stage pipeline. Inference time is fast and can run on a NVIDIA TITAN XP GPU at 1 HZ for eight or less objects present. While only trained on simulated data, CARTO transfers to real-world object instances. Code and evaluation data is available at: carto. cs. uni - freiburg. de
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- PARIS: Part-level Reconstruction and Motion Analysis for Articulated ObjectsJiayi Liu, Ali Mahdavi-Amiri, Manolis SavvaICCV 2023 · 103 citations
- Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated ObjectsChuanruo Ning, Ruihai Wu, Haoran Lu, Kaichun Mo et al.NeurIPS 2023 · 64 citations
- NeO 360: Neural Fields for Sparse View Synthesis of Outdoor ScenesMuhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Guizilini et al.ICCV 2023 · 63 citations
- Articulate your NeRF: Unsupervised articulated object modeling via conditional view synthesisJianning Deng, Kartic Subr, Hakan BilenNeurIPS 2024 · 26 citations
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelZhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu et al.NeurIPS 2025 · 24 citations
Builds on16
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and PoseShih-Yang Su, Frank Yu, Michael Zollhöfer, Helge RhodinNeurIPS 2021 · 316 citations
- A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationJiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille et al.ICCV 2021 · 138 citations
- CAPTRA: CAtegory-level Pose Tracking for Rigid and Articulated Objects from Point CloudsYijia Weng, He Wang, Qiang Zhou, Yuzhe Qin et al.ICCV 2021 · 119 citations
- Ditto: Building Digital Twins of Articulated Objects from InteractionZhenyu Jiang, Cheng-Chun Hsu, Yuke ZhuCVPR 2022 · 77 citations
- AKB-48: A Real-World Articulated Object Knowledge BaseLiu Liu, Wenqiang Xu, Haoyuan Fu, Sucheng Qian et al.CVPR 2022 · 64 citations
Related papers
- FroDO: From Detections to 3D ObjectsMartin Rünz, Kejie Li, Meng Tang, Lingni Ma et al.CVPR 2020
- ART: Articulated Reconstruction TransformerZizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins et al.CVPR 2026 · 12 citations
- Detection Based Part-level Articulated Object Reconstruction from Single RGBD ImageYuki Kawana, Tatsuya HaradaNeurIPS 2023 · 20 citations
- Multi-Path Learning for Object Pose Estimation Across DomainsMartin Sundermeyer, Maximilian Durner, En Yen Puang, Zoltan-Csaba Marton et al.CVPR 2020
- R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyLi Zhang, Haonan Jiang, Yukang Huo, Yan Zhong et al.AAAI 2025 · 6 citations
