NCHO: Unsupervised Learning for Neural 3D Composition of Humans and Objects
Taeksoo Kim, Shunsuke Saito, Hanbyul Joo
Abstract
Deep generative models have been recently extended to synthesizing 3D digital humans. However, previous approaches treat clothed humans as a single chunk of geometry without considering the compositionality of clothing and accessories. As a result, individual items cannot be naturally composed into novel identities, leading to limited expressiveness and controllability of generative 3D avatars. While several methods attempt to address this by leveraging synthetic data, the interaction between humans and objects is not authentic due to the domain gap, and manual asset creation is difficult to scale for a wide variety of objects. In this work, we present a novel framework for learning a compositional generative model of humans and objects (backpacks, coats, scarves, and more) from real-world 3D scans. Our compositional model is interaction-aware, meaning the spatial relationship between humans and objects, and the mutual shape change by physical contact is fully incorporated. The key challenge is that, since humans and objects are in contact, their 3D scans are merged into a single piece. To decompose them without manual annotations, we propose to leverage two sets of 3D scans of a single person with and without objects. Our approach learns to decompose objects and naturally compose them back into a generative human model in an unsupervised manner. Despite our simple setup requiring only the capture of a single subject with objects, our experiments demonstrate the strong generalization of our model by enabling the natural composition of objects to diverse identities in various poses and the composition of multiple objects, which is unseen in training data. The project page is available at https://taeksuu.github.io/ncho.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- Template Free Reconstruction of Human-object Interaction with Procedural Interaction GenerationXianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, Gerard Pons-MollCVPR 2024 · 6 citations
- HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsByungjun Kim, Shunsuke Saito, Giljoo Nam, Tomas Simon et al.ICCV 2025 · 2 citations
Builds on35
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Implicit Geometric Regularization for Learning ShapesAmos Gropp, Lior Yariv, Niv Haim, Matan Atzmon et al.ICML 2020 · 1,001 citations
Related papers
- AG3D: Learning to Generate 3D Avatars from 2D Image CollectionsZijian Dong, Xu Chen, Jinlong Yang, Michael J. Black et al.ICCV 2023 · 76 citations
- Interact-Custom: Customized Human Object Interaction Image GenerationZhu Xu, Zhaowen Wang, Yuxin Peng, Yang LiuACM MM 2025 · 1 citation
- gDNA: Towards Generative Detailed Neural AvatarsXu Chen, Tianjian Jiang, Jie Song, Jinlong Yang et al.CVPR 2022 · 71 citations
- SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human MeshesSoubhik Sanyal, Partha Ghosh, Jinlong Yang, Michael J. Black et al.CVPR 2024 · 2 citations
- Self-Supervised Collision Handling via Generative 3D Garment Models for Virtual Try-OnIgor Santesteban, Nils Thuerey, Miguel A. Otaduy, Dan CasasCVPR 2021
