Detection Based Part-level Articulated Object Reconstruction from Single RGBD Image
Yuki Kawana, Tatsuya Harada
摘要
We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from previous works that rely on learning instance-level latent space, focusing on man-made articulated objects with predefined part counts. Instead, we propose a novel alternative approach that employs part-level representation, representing instances as combinations of detected parts. While our detect-then-group approach effectively handles instances with diverse part structures and various part counts, it faces issues of false positives, varying part sizes and scales, and an increasing model size due to end-to-end training. To address these challenges, we propose 1) test-time kinematics-aware part fusion to improve detection performance while suppressing false positives, 2) anisotropic scale normalization for part shape learning to accommodate various part sizes and scales, and 3) a balancing strategy for cross-refinement between feature space and output space to improve part detection while maintaining model size. Evaluation on both synthetic and real data demonstrates that our method successfully reconstructs variously structured multiple instances that previous works cannot handle, and outperforms prior works in shape reconstruction and kinematics estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Articulate your NeRF: Unsupervised articulated object modeling via conditional view synthesisJianning Deng, Kartic Subr, Hakan BilenNeurIPS 2024 · 被引用 26 次
- FreeArtGS: Articulated Gaussian Splatting Under Free-moving ScenarioHang Dai, Hongwei Fan, Han Zhang, Duojin Wu 等CVPR 2026 · 被引用 3 次
- Monomobility: Zero-Shot 3D Mobility Analysis From Monocular VideosHongyi Zhou, Yulan Guo, Xiaogang Wang, Kai XuICCV 2025 · 被引用 3 次
- ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility ProposalsXuelu Li, Zhaonan Wang, Xiaogang Wang, Lei Wu 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper22
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 被引用 602 次
- Nerfstudio: A Modular Framework for Neural Radiance Field DevelopmentMatthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li 等SIGGRAPH 2023 · 被引用 592 次
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera 等ICCV 2019 · 被引用 504 次
- Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. BlackICCV 2019 · 被引用 183 次
相关 Paper
- CARTO: Category and Joint Agnostic Reconstruction of ARTiculated ObjectsNick Heppert, Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu 等CVPR 2023
- From Points to Multi-Object 3D ReconstructionFrancis Engelmann, Konstantinos Rematas, Bastian Leibe, Vittorio FerrariCVPR 2021
- ART: Articulated Reconstruction TransformerZizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins 等CVPR 2026 · 被引用 12 次
- Learning Canonical Shape Space for Category-Level 6D Object Pose and Size EstimationDengsheng Chen, Jun Li, Zheng Wang, Kai XuCVPR 2020
- Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) EquivarianceXueyi Liu, Ji Zhang, Ruizhen Hu, Haibin Huang 等ICLR 2023 · 被引用 3 次
