ScanBot: Autonomous Reconstruction via Deep Reinforcement Learning
Hezhi Cao, Xi Xia, Guan Wu, Ruizhen Hu, Ligang Liu
Abstract
Autoscanning of an unknown environment is the key to many AR/VR and robotic applications. However, autonomous reconstruction with both high efficiency and quality remains a challenging problem. In this work, we propose a reconstruction-oriented autoscanning approach, called ScanBot, which utilizes hierarchical deep reinforcement learning techniques for global region-of-interest (ROI) planning to improve the scanning efficiency and local next-best-view (NBV) planning to enhance the reconstruction quality. Given the partially reconstructed scene, the global policy designates an ROI with insufficient exploration or reconstruction. The local policy is then applied to refine the reconstruction quality of objects in this region by planning and scanning a series of NBVs. A novel mixed 2D-3D representation is designed for these policies, where a 2D quality map with tailored quality channels encoding the scanning progress is consumed by the global policy, and a coarse-to-fine 3D volumetric representation that embodies both local environment and object completeness is fed to the local policy. These two policies iterate until the whole scene has been completely explored and scanned. To speed up the learning of complex environmental dynamics and enhance the agent's memory for spatial-temporal inference, we further introduce two novel auxiliary learning tasks to guide the training of our global policy. Thorough evaluations and comparisons are carried out to show the feasibility of our proposed approach and its advantages over previous methods. Code and data are available at https://github.com/HezhiCao/Scanbot.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7a89093d-0e09-4273-8f4c-3fee4e2bb54fRelated papers
- Active3D: Active High-Fidelity 3D Reconstruction via Multi-Level Uncertainty QuantificationYan Li, Yingzhao Li, Gim Hee LeeAAAI 2026
- Appearance-aware Multi-view SVBRDF Reconstruction via Deep Reinforcement LearningPengfei Zhu, Jie Guo, Yifan Liu, Qi Sun et al.SIGGRAPH 2025 · 1 citation
- CSO: Constraint-Guided Space Optimization for Active Scene MappingXuefeng Yin, Chenyang Zhu, Shanglai Qu, Yuqi Li et al.ACM MM 2024
- Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via DiscretisationStephen James, Kentaro Wada, Tristan Laidlow, Andrew J. DavisonCVPR 2022 · 65 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
