DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
Ruiqi Wu, Xinjie Wang, Liu Liu, Chun-Le Guo, Jiaxiong Qiu, Chongyi Li, Lichao Huang, Zhizhong Su, Ming-Ming Cheng
Abstract
We present DIPO, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for data collection, but at the same time provides important motion information, which is a reliable guide for predicting kinematic relationships between parts. Specifically, we propose a dual-image diffusion model that captures relationships between the image pair to generate part layouts and joint parameters. In addition, we introduce a Chain-of-Thought (CoT) based graph reasoner that explicitly infers part connectivity relationships. To further improve robustness and generalization on complex articulated objects, we develop a fully automated dataset expansion pipeline, name LEGO-Art, that enriches the diversity and complexity of PartNet-Mobility dataset. We propose PM-X, a large-scale dataset of complex articulated 3D objects, accompanied by rendered images, URDF annotations, and textual descriptions. Extensive experiments demonstrate that DIPO significantly outperforms existing baselines in both the resting state and the articulated state, while the proposed PM-X dataset further enhances generalization to diverse and structurally complex articulated objects. Our code and dataset will be released to the community upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- ArtLLM: Generating Articulated Assets via 3D LLMPenghao Wang, Siyuan Xie, Hongyu Yan, Xianghui Yang et al.CVPR 2026 · 7 citations
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM KnowledgeYumeng He, Ying Jiang, Jiayin Lu, Yin Yang et al.CVPR 2026 · 6 citations
- Artiverse: A Diverse and Physically Grounded Dataset for Articulated ObjectsDenys Iliash, Jiayi Liu, Egor Fokin, Qirui Wu et al.CVPR 2026 · 4 citations
- Repurposing 3D Generative Model for Autoregressive Layout GenerationHaoran Feng, Yifan Niu, Zehuan Huang, Yangtian Sun et al.CVPR 2026 · 3 citations
- ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility ProposalsXuelu Li, Zhaonan Wang, Xiaogang Wang, Lei Wu et al.CVPR 2026 · 1 citation
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- NAP: Neural 3D Articulated Object PriorJiahui Lei, Congyue Deng, William B. Shen, Leonidas J. Guibas et al.NeurIPS 2023 · 53 citations
- SINGAPO: Single Image Controlled Generation of Articulated Parts in ObjectsJiayi Liu, Denys Iliash, Angel X. Chang, Manolis Savva et al.ICLR 2025
- BrickNet: Graph-Backed Generative Brick AssemblyPeter Kulits, Cordelia SchmidCVPR 2026 · 5 citations
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht et al.CVPR 2026 · 22 citations
- DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-SpacesLi Zhang, Mingyu Mei, Ailing Wang, Xianhui Meng et al.CVPR 2026 · 2 citations
