SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects
Jiayi Liu, Denys Iliash, Angel X. Chang, Manolis Savva, Ali Mahdavi Amiri
摘要
We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation process. These limitations hinder the scalability and practicality for articulated object modeling. In this work, we propose a method to generate articulated objects from a single image. Observing the object in resting state from an arbitrary view, our method generates an articulated object that is visually consistent with the input image. To capture the ambiguity in part shape and motion posed by a single view of the object, we design a diffusion model that learns the plausible variations of objects in terms of geometry and kinematics. To tackle the complexity of generating structured data with attributes in multiple domains, we design a pipeline that produces articulated objects from high-level structure to geometric details in a coarse-to-fine manner, where we use a part connectivity graph and part abstraction as proxies. Our experiments show that our method outperforms the state-of-theart in articulated object creation by a large margin in terms of the generated object realism, resemblance to the input image, and reconstruction quality. RELATED WORK Generation of structured data. Our task is closely related to the generation of structured data (Chaudhuri et al., 2020) . The generation of 3D shapes with semantic parts is a widely studied problem with the main goal of modeling objects with geometric details and semantic grouping at the part level. Prior work either synthesizes objects in voxels with semantic labels (Wang et al., 2018; Li et al., 2020; Wu et al., 2020) or further considers the spatial structure, such as symmetry and support relationship, by jointly learning in the latent space (Wu et al., 2019) or explicitly modeling
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht 等CVPR 2026 · 被引用 22 次
- Dexterous World ModelsByungjun Kim, Taeksoo Kim, Junyoung Lee, Hanbyul JooCVPR 2026 · 被引用 17 次
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao 等CVPR 2026 · 被引用 8 次
- Guiding Diffusion-Based Articulated Object Generation by Partial Point Cloud Alignment and Physical Plausibility ConstraintsJens U. Kreber, Joerg StuecklerICCV 2025 · 被引用 7 次
- PDGS: Part-Level Decoupling and Continuous Deformation of Articulated Objects via Gaussian SplattingHaowen Wang, Xiaoping Yuan, Zhao Jin, Zhen Zhao 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper41
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
相关 Paper
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM KnowledgeYumeng He, Ying Jiang, Jiayin Lu, Yin Yang 等CVPR 2026 · 被引用 6 次
- Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsYukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu 等ACM MM 2024 · 被引用 20 次
- Expressive Talking Human from Single-Image with Imperfect PriorsJun Xiang, Yudong Guo, Leipeng Hu, Boyang Guo 等ICCV 2025 · 被引用 3 次
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang 等NeurIPS 2023 · 被引用 65 次
- Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorJunshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang 等ICCV 2023 · 被引用 405 次
