Learning Object Context for Novel-view Scene Layout Generation
Xiaotian Qiao, Gerhard P. Hancke, Rynson W. H. Lau
Abstract
Novel-view prediction of a scene has many applications. Existing works mainly focus on generating novel-view images via pixel-wise prediction in the image space, often resulting in severe ghosting and blurry artifacts. In this paper, we make the first attempt to explore novel-view prediction in the layout space, and introduce the new problem of novel-view scene layout generation. Given a single scene layout and the camera transformation as inputs, our goal is to generate a plausible scene layout for a specified viewpoint. Such a problem is challenging as it involves accurate understanding of the 3D geometry and semantics of the scene from as little as a single 2D scene layout. To tackle this challenging problem, we propose a deep model to capture contextualized object representation by explicitly modeling the object context transformation in the scene. The contextualized object representation is essential in generating geometrically and semantically consistent scene layouts of different views. Experiments show that our model outperforms several strong baselines on many indoor and outdoor scenes, both qualitatively and quantitatively. We also show that our model enables a wide range of applications, including novel-view image synthesis, novel-view image editing, and amodal object estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45a33dec-3233-4299-8cd5-ed50e67d393dCited by top-tier papers3
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- Preserving Structural Consistency in Arbitrary Artist and Artwork Style TransferJingyu Wu, Lefan Hou, Zejian Li, Jun Liao et al.AAAI 2023 · 6 citations
- Painting 3D Nature in 2D: View Synthesis of Natural Scenes from a Single Semantic MaskShangzhan Zhang, Sida Peng, Tianrun Chen, Linzhan Mou et al.CVPR 2023
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- iMAP: Implicit Mapping and Positioning in Real-TimeEdgar Sucar, Shikun Liu, Joseph Ortiz, Andrew J. DavisonICCV 2021 · 834 citations
- In-Place Scene Labelling and Understanding with Implicit Scene RepresentationShuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. DavisonICCV 2021 · 551 citations
- Extreme View SynthesisInchang Choi, Orazio Gallo, Alejandro J. Troccoli, Min H. Kim et al.ICCV 2019 · 207 citations
- LayoutVAE: Stochastic Scene Layout Generation From a Label SetAkash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal et al.ICCV 2019 · 194 citations
Related papers
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng et al.CVPR 2021
- Generative View Synthesis: From Single-view Semantics to Novel-view ImagesTewodros Amberbir Habtegebrial, Varun Jampani, Orazio Gallo, Didier StrickerNeurIPS 2020 · 20 citations
- Layout-Guided Novel View Synthesis From a Single Indoor PanoramaJiale Xu, Jia Zheng, Yanyu Xu, Rui Tang et al.CVPR 2021
- LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly DetectionZhichao Zeng, Jiasheng Zhang, Jiyun Sun, Jiangtao Cui et al.CVPR 2026
- PanoContext-Former: Panoramic Total Scene Understanding with a TransformerYuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong et al.CVPR 2024
