RelScene: A Benchmark and baseline for Spatial Relations in text-driven 3D Scene Generation
Zhaoda Ye, Xinhan Zheng, Yang Liu, Yuxin Peng
Abstract
Text-driven 3D indoor scene generation aims to automatically generate and arrange the objects, which form a 3D scene that accurately captures the semantics detailed in the given text description. Recent works have shown the potential to generate 3D scenes guided by specific object categories and room layouts but lack a robust mechanism to maintain consistent spatial relationships in alignment with the provided text description during the 3D scene generation. Besides, the annotations of the object and relationships of the 3D scenes are usually time- and cost-consuming, which are not easily obtained for the model training. Thus, in this paper, we conduct a dataset and benchmark for assessing spatial relations in text-driven 3D scene generation, which contains a comprehensive collection of 3D scenes, including textual descriptions, annotating object spatial relations, and providing both template and free-form natural language descriptions. We also provide a pseudo description feature generation method to address the 3D scenes without language annotations. We design an aligned latent space for spatial relation in 3D scenes and text description, in which we can sample the features according to the spatial relation for the few-shot learning. We also propose new metrics to investigate the ability of the approach to generate correct spatial relationships among objects.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9c9de00e-896f-4df5-bf3d-5f63a1cecf6cRelated papers
- M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D GenerationYiheng Zhang, Zhuojiang Cai, Mingdao Wang, Meitong Guo et al.CVPR 2026 · 5 citations
- InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph PriorChenguo Lin, Yadong MuICLR 2024 · 94 citations
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 19 citations
- Text-to-Scene with Large Reasoning ModelsFrédéric Berdoz, Luca A. Lanzendörfer, Nick Tuninga, Roger WattenhoferAAAI 2026 · 2 citations
