Learning 3D Object Spatial Relationships From Pre-Trained 2D Diffusion Models
Sangwon Baik, Hyeonwoo Kim, Hanbyul Joo
Abstract
We present a method for learning 3D spatial relationships between object pairs, referred to as object-object spatial relationships (OOR), by leveraging synthetically generated 3D samples from pre-trained 2D diffusion models. We hypothesize that images synthesized by diffusion models inherently capture realistic OOR cues, enabling efficient collection of a 3D dataset to learn OOR for various unbounded object categories. Our approach synthesizes diverse images that capture plausible OOR cues, which we then uplift into 3D samples. Leveraging our diverse collection of 3D samples for the object pairs, we train a score-based OOR diffusion model to learn the distribution of their relative spatial relationships. Additionally, we extend our pairwise OOR to multi-object OOR by enforcing consistency across pairwise relations and preventing object collisions. Extensive experiments demonstrate the robustness of our method across various object-object spatial relationships, along with its applicability to 3D scene arrangement tasks and human motion synthesis using our OOR diffusion model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0c8774e-071d-4a17-b62f-8d714ad3a021Cited by top-tier papers4
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 7 citations
- Learning to Generate Human-Human-Object Interactions from Textual DescriptionsJeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee et al.NeurIPS 2025 · 3 citations
- DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-Trained Video Diffusion ModelsHyeonwoo Kim, Sangwon Baik, Hanbyul JooICCV 2025 · 1 citation
- Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric ConstraintsRotem Gatenyo, Ohad FriedCVPR 2026
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
Related papers
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 19 citations
- HOLODIFFUSION: Training a 3D Diffusion Model Using 2D ImagesAnimesh Karnewar, Andrea Vedaldi, David Novotný, Niloy J. MitraCVPR 2023
- CAD : Photorealistic 3D Generation via Adversarial DistillationZiyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu et al.CVPR 2024 · 3 citations
- CG-HOI: Contact-Guided 3D Human-Object Interaction GenerationChristian Diller, Angela DaiCVPR 2024
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D DiffusionHongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu et al.CVPR 2026 · 2 citations
