Monte Carlo Scene Search for 3D Scene Understanding
Shreyas Hampali, Sinisa Stekovic, Sayan Deb Sarkar, Chetan Srinivasa Kumar, Friedrich Fraundorfer, Vincent Lepetit
Abstract
Abstract We explore how a general AI algorithm can be used for 3D scene understanding to reduce the need for training data. More exactly, we propose a modification of the Monte Carlo Tree Search (MCTS) algorithm to retrieve objects and room layouts from noisy RGB-D scans. While MCTS was developed as a game-playing algorithm, we show it can also be used for complex perception problems. Our adapted MCTS algorithm has few easy-to-tune hyperparameters and can optimise general losses. We use it to optimise the posterior prob-ability of objects and room layout hypotheses given the RGB-D data. This results in an analysis-by-synthesis approach that explores the solution space by rendering the current solution and comparing it to the RGB-D observations. To perform this exploration even more efficiently, we propose simple changes to the standard MCTS' tree construction and exploration policy. We demonstrate our approach on the ScanNet dataset. Our method often retrieves configurations that are better than some manual annotations, especially on layouts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- MonteFloor: Extending MCTS for Reconstructing Accurate Large-Scale Floor PlansSinisa Stekovic, Mahdi Rad, Friedrich Fraundorfer, Vincent LepetitICCV 2021 · 43 citations
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao et al.NeurIPS 2025 · 26 citations
- Convex Decomposition of Indoor ScenesVaibhav Vavilala, David A. ForsythICCV 2023 · 11 citations
- DDIT: Semantic Scene Completion via Deformable Deep Implicit TemplatesHaoang Li, Jinhu Dong, Binghui Wen, Ming Gao et al.ICCV 2023 · 8 citations
- DeepSPF: Spherical SO(3)-Equivariant Patches for Scan-to-CAD EstimationDriton Salihu, Adam Misik, Yuankai Wu, Constantin Patsch et al.ICLR 2024 · 2 citations
Builds on7
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseYixin Chen, Siyuan Huang, Tao Yuan, Yixin Zhu et al.ICCV 2019 · 130 citations
- Joint Embedding of 3D Scan and CAD ObjectsManuel Dahnert, Angela Dai, Leonidas J. Guibas, Matthias NießnerICCV 2019 · 36 citations
- MSeg: A Composite Dataset for Multi-Domain Semantic SegmentationJohn Lambert, Zhuang Liu, Ozan Sener, James Hays et al.CVPR 2020
Related papers
- Reinforcement Learning and Data-Generation for Syntax-Guided SynthesisJulian Parsert, Elizabeth PolgreenAAAI 2024 · 7 citations
- SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingDengwei Zhao, Shikui Tu, Lei XuNeurIPS 2024 · 4 citations
- RandomRooms: Unsupervised Pre-training from Synthetic Shapes and Randomized Layouts for 3D Object DetectionYongming Rao, Benlin Liu, Yi Wei, Jiwen Lu et al.ICCV 2021 · 58 citations
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- Shape Anchor Guided Holistic Indoor Scene UnderstandingMingyue Dong, Linxi Huan, Hanjiang Xiong, Shuhan Shen et al.ICCV 2023 · 5 citations
