CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
Kaixin Yao, Longwen Zhang, Xinhao Yan, Yan Zeng, Qixuan Zhang, Lan Xu, Wei Yang, Jiayuan Gu, Jingyi Yu
Abstract
Recovering high-quality 3D scenes from a single RGB image is a challenging task in computer graphics. Current methods often struggle with domain-specific limitations or low-quality object generation. To address these, we propose CAST (Component-Aligned 3D Scene Reconstruction from a Single RGB Image), a novel method for 3D scene reconstruction. CAST starts by extracting object-level 2D segmentation and relative depth information from the input image, followed by using a GPT-based model to analyze inter-object spatial relations. This enables understanding of how objects relate to each other within the scene, ensuring more coherent reconstruction. CAST then employs an occlusion-aware large-scale 3D generation model to independently generate each object's full geometry, using Masked Auto Encoder (MAE) and point cloud conditioning to mitigate the effects of occlusions and partial object information, ensuring accurate alignment with the source image's geometry and texture. To align each object with the scene, the alignment generation model computes the necessary transformations, allowing the generated meshes to be accurately placed and integrated into the scene's point cloud. Finally, CAST applies a physics-aware correction mechanism, which leverages a fine-grained relation graph to generate a constraint graph. This graph guides the optimization of object poses, ensuring physical consistency and spatial coherence. By utilizing Signed Distance Fields (SDF), the model effectively addresses issues such as occlusions, object penetration, and floating objects, ensuring that the generated scene accurately reflects real-world physical interactions. Experimental results demonstrate that CAST significantly improves the quality of single-image 3D scene reconstruction, offering enhanced realism and accuracy in scene understanding and reconstruction tasks. CAST has practical applications in virtual content creation, such as immersive game environments and film production, where real-world setups can be seamlessly integrated into virtual landscapes. Additionally, CAST can be leveraged in robotics, enabling efficient real-to-simulation workflows and providing realistic, scalable simulation environments for robotic systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14d44e58-de00-4d5e-91aa-7f564eddefb7Cited by top-tier papers31
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang et al.CVPR 2026 · 280 citations
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao et al.NeurIPS 2025 · 26 citations
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoningJinkun Hao, Naifu Liang, Zhen Luo, Xudong Xu et al.NeurIPS 2025 · 21 citations
- ShapeR: Robust Conditional 3D Shape Generation from Casual CapturesYawar Siddiqui, Duncan P. Frost, Samir Aroudj, Armen Avetisyan et al.CVPR 2026 · 19 citations
Builds on43
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
Related papers
- PixARMesh: Autoregressive Mesh-Native Single-View Scene ReconstructionXiang Zhang, Sohyun Yoo, Hongrui Wu, Chuan Li et al.CVPR 2026 · 2 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- SimRecon: SimReady Compositional Scene Reconstruction from Real VideosChong Xia, Kai Zhu, Zizhuo Wang, Fangfu Liu et al.CVPR 2026 · 11 citations
- HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene GenerationBin Zang, Wenting Zheng, Xiaoliang Luo, Zhiyuan Fang et al.ICML 2026
- CASAGPT: Cuboid Arrangement and Scene Assembly for Interior DesignWeitao Feng, Hang Zhou, Jing Liao, Li Cheng et al.CVPR 2025
