SKDream: Controllable Multi-view and 3D Generation with Arbitrary Skeletons
Yuanyou Xu, Zongxin Yang, Yi Yang
Abstract
Controllable generation has achieved substantial progress in both 2D and 3D domains, yet current conditional generation methods still face limitations in describing detailed shape structures. Skeletons can effectively represent and describe object anatomy and pose. Unfortunately, past studies are often limited to human skeletons. In this work, we generalize skeletal conditioned generation to arbitrary structures. First, we design a reliable mesh skeletonization pipeline to generate a large-scale mesh-skeleton paired dataset. Based on the dataset, a multi-view and 3D generation pipeline is built. We propose to represent 3D skeletons by Coordinate Color Encoding as 2D conditional images. A Skeletal Correlation Module is designed to extract global skeletal features for condition injection. After multi-view images are generated, 3D assets can be obtained by incorporating a large reconstruction model, followed by a UV texture refinement stage. As a result, our method achieves instant generation of multi-view and 3D contents that are aligned with given skeletons. The proposed techniques largely improve the object-skeleton alignment and generation quality. Project page at https://skdream3d.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3a7518a-aeda-4f31-aae4-305845fb4309Cited by top-tier papers3
- Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion modelsRuisi Zhao, Haoren Zheng, Zongxin Yang, Hehe Fan et al.ICLR 2026 · 2 citations
- PoseMaster: A Unified 3D Native Framework for Stylized Pose GenerationHongyu Yan, Kunming Luo, Weiyu Li, Kaiyi Zhang et al.CVPR 2026 · 1 citation
- Insert Anything: Image Insertion via In-Context Editing in DiTWensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng et al.AAAI 2026 · 1 citation
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Animator-Centric Skeleton Generation on Objects with Fine-Grained DetailsMingze Sun, Cheng Zeng, Jiansong Pei, Junhao Chen et al.CVPR 2026 · 10 citations
- ARMO: Autoregressive Rigging for Multi-Category ObjectsMingze Sun, Shiwei Mao, Keyi Chen, Yurun Chen et al.ICCV 2025 · 3 citations
- Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without TrainingHexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo et al.CVPR 2026
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesYatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng et al.ICCV 2025 · 2 citations
- Structured 3D Latents for Scalable and Versatile 3D GenerationJianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng et al.CVPR 2025
