Consistent Depth Prediction for Transparent Object Reconstruction from RGB-D Camera
Yuxiang Cai, Yifan Zhu, Haiwei Zhang, Bo Ren
Abstract
Transparent objects are commonly seen in indoor scenes but are hard to estimate. Currently, commercial depth cameras face difficulties in estimating the depth of transparent objects due to the light reflection and refraction on their surface. As a result, they tend to make a noisy and incorrect depth value for transparent objects. These incorrect depth data make the traditional RGB-D SLAM method fails in reconstructing the scenes that contain transparent objects. An exact depth value of the transparent object is required to restore in advance and it is essential that the depth value of the transparent object must keep consistent in different views, or the reconstruction result will be distorted. Previous depth prediction methods of transparent objects can restore these missing depth values but none of them can provide a good result in reconstruction due to the inconsistency prediction. In this work, we propose a real-time reconstruction method using a novel stereo-based depth prediction network to keep the consistency of depth prediction in a sequence of images. Because there is no video dataset about transparent objects currently to train our model, we construct a synthetic RGB-D video dataset with different transparent objects. Moreover, to test generalization capability, we capture video from real scenes using the RealSense D435i RGB-D camera. We compare the metrics on our dataset and SLAM reconstruction results in both synthetic scenes and real scenes with the previous methods. Experiments show our significant improvement in accuracy on depth prediction and scene reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d24c57df-07e0-4373-9d78-3d3adf4f701fCited by top-tier papers1
Ask how each one uses itBuilds on6
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe et al.ICCV 2021 · 246 citations
- Interpolated Convolutional Networks for 3D Point Cloud UnderstandingJiageng Mao, Xiaogang Wang, Hongsheng LiICCV 2019 · 241 citations
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- Transfusion: A Novel SLAM Method Focused on Transparent ObjectsYifan Zhu, Jiaxiong Qiu, Bo RenICCV 2021 · 20 citations
- RGB-D Local Implicit Function for Depth Completion of Transparent ObjectsLuyang Zhu, Arsalan Mousavian, Yu Xiang, Hammad Mazhar et al.CVPR 2021
Related papers
- Through the Looking Glass: Neural 3D Reconstruction of Transparent ShapesZhengqin Li, Yu-Ying Yeh, Manmohan ChandrakerCVPR 2020
- 2D Gaussian Splatting-Based Sparse-View Transparent Object Depth Reconstruction Via Physics Simulation for Scene UpdateJeongyun Kim, Seunghoon Jeong, Giseop Kim, Myung-Hwan Jeon et al.ICCV 2025 · 1 citation
- RTG-SLAM: Real-time 3D Reconstruction at Scale using Gaussian SplattingZhexi Peng, Tianjia Shao, Yong Liu, Jingke Zhou et al.SIGGRAPH 2024 · 96 citations
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova et al.CVPR 2023
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular VideosShi Chen, Erik Sandström, Sandro Lombardi, Siyuan Li et al.NeurIPS 2025 · 1 citation
