3D Software Synthesis Driven by Constraint-Expressive Intermediate Representation
Shuqing Li, Anson Y. Lam, Yun Peng, Wenxuan Wang, Michael R. Lyu
摘要
Graphical user interface (UI) software has undergone a fundamental transformation from traditional 2D interfaces to spatial 3D environments. While existing work has made remarkable success in 2D software generation, 3D software generation still remains underexplored. Current methods for 3D software generation usually generate 3D environment as a whole and cannot modify specific elements. Furthermore, these methods struggle to handle the complex spatial and semantic constraints inherent in the real world.
To address these challenges, we present Scenethesis, a novel requirement-sensitive 3D software synthesis approach that maintains formal traceability between user specifications and generated 3D software. Scenethesis is built upon ScenethesisLang, a domain-specific language that serves as a granular constraintaware intermediate representation to bridge natural language requirements and 3D software. It serves both as a comprehensive scene description language enabling fine-grained modification of 3D software elements and as a formal constraint-expressive specification language capable of expressing complex spatial constraints. By decomposing 3D software synthesis into stages, Scenethesis enables independent verification, targeted modification, and systematic constraint satisfaction. Our evaluation demonstrates that Scenethesis accurately captures over 80% of user requirements and satisfies more than 90% of hard constraints while handling over 100 constraints simultaneously. Furthermore, Scenethesis achieves a 42.8% improvement in BLIP-2 visual evaluation scores compared to the state-of-the-art method, establishing its effectiveness in generating high-quality 3D software that faithfully adheres to complex user requirements. Source code and supplemental materials are publicly available at https://sites.google.com/view/3d-software-synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis 等NeurIPS 2021 · 被引用 293 次
相关 Paper
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding 等ICLR 2026 · 被引用 74 次
- SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionYueming Zhao, Hongyu Yang, Di HuangAAAI 2026
- SceneX: Procedural Controllable Large-Scale Scene GenerationMengqi Zhou, Yuxi Wang, Jun Hou, Shougao Zhang 等AAAI 2025 · 被引用 22 次
- Scenepainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentChong Xia, Shengjun Zhang, Fangfu Liu, Chang Liu 等ICCV 2025 · 被引用 2 次
- In Situ 3D Scene Synthesis for Ubiquitous Embodied InterfacesHaiyan Jiang, Leiyu Song, Dongdong Weng, Zhe Sun 等ACM MM 2024 · 被引用 3 次
