Generative Proxemics: A Prior for 3D Social Interaction from Images
Lea Müller, Vickie Ye, Georgios Pavlakos, Michael J. Black, Angjoo Kanazawa
Abstract
Social interaction is a fundamental aspect of human behavior and communication. The way individuals position themselves in relation to others, also known as proxemics, conveys social cues and affects the dynamics of social interaction. Reconstructing such interaction from images presents challenges because of mutual occlusion and the limited availability of large training datasets. To address this, we present a novel approach that learns a prior over the 3D proxemics two people in close social interaction and demonstrate its use for single-view 3D reconstruction. We start by creating 3D training data of interacting people using image datasets with contact annotations. We then model the proxemics using a novel denoising diffusion model called BUDDI that learns the joint distribution over the poses of two people in close social interaction. Sampling from our generative proxemics model produces realistic 3D human interactions, which we validate through a perceptual study. We use BUDDI in reconstructing two people in close proximity from an image without any contact annotation via an optimization approach that uses the diffusion model as a prior. Our approach recovers accurate 3D social interactions from noisy initial estimates, outperforming state-of-the-art methods. Our code, data, and model are available at: muelea. github.io/buddi.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Iterative Motion Editing with Natural LanguagePurvi Goel, Kuan-Chieh Wang, C. Karen Liu, Kayvon FatahalianSIGGRAPH 2024 · 22 citations
- MultiPhys: Multi-Person Physics-Aware 3D Motion EstimationNicolas Ugrinovic, Boxiao Pan, Georgios Pavlakos, Despoina Paschalidou et al.CVPR 2024 · 5 citations
- DPoser-X: Diffusion Model as Robust 3D Whole-Body Human Pose PriorJunzhe Lu, Jing Lin, Hongkun Dou, Ailing Zeng et al.ICCV 2025 · 5 citations
- DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingJungbin Cho, Junwan Kim, Jisoo Kim, Minseo Kim et al.ICCV 2025 · 4 citations
Builds on38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan et al.CVPR 2024
- ContactGen: Contact-Guided Interactive 3D Human Generation for PartnersDongjun Gu, Jaehyeok Shim, Jaehoon Jang, Changwoo Kang et al.AAAI 2024 · 5 citations
- Template Free Reconstruction of Human-object Interaction with Procedural Interaction GenerationXianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, Gerard Pons-MollCVPR 2024 · 6 citations
- Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric ViewsSiwei Zhang, Qianli Ma, Yan Zhang, Sadegh Aliakbarian et al.ICCV 2023 · 34 citations
- Ponimator: Unfolding Interactive Pose for Versatile Human-Human Interaction AnimationShaowei Liu, Chuan Guo, Bing Zhou, Jian WangICCV 2025 · 2 citations
