Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth Estimation
Hongyu Wen, Yiming Zuo, Venkat Subramanian, Patrick Chen, Jia Deng
Abstract
Transparent objects are common in daily life, and understanding their multi-layer depth information -- perceiving both the transparent surface and the objects behind it -- is crucial for real-world applications that interact with transparent materials. In this paper, we introduce LayeredDepth, the first dataset with multi-layer depth annotations, including a real-world benchmark and a synthetic data generator, to support the task of multi-layer depth estimation. Our real-world benchmark consists of 1,500 images from diverse scenes, and evaluating state-of-the-art depth estimation methods on it reveals that they struggle with transparent objects. The synthetic data generator is fully procedural and capable of providing training data for this task with an unlimited variety of objects and scene compositions. Using this generator, we create a synthetic dataset with 15,300 images. Baseline models training solely on this synthetic dataset produce good cross-domain multi-layer depth estimation. Fine-tuning state-of-the-art single-layer depth models on it substantially improves their performance on transparent objects, with quadruplet accuracy on our benchmark increased from 55.14% to 75.20%. All images and validation annotations are available under CC0 at https://layereddepth.cs.princeton.edu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Sharp Monocular View Synthesis in Less Than a SecondLars Mescheder, Wei Dong, Shiwei Li, Xuyang Bai et al.ICLR 2026 · 25 citations
- DepthFocus: Controllable Depth Estimation for See-Through Scenesjunhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min et al.CVPR 2026 · 4 citations
- SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined GroupingHongyu Wen, Jia DengCVPR 2026
Builds on19
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
- UniDepth: Universal Monocular Metric Depth EstimationLuigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segù et al.CVPR 2024 · 122 citations
Related papers
- Towards Multimodal Depth Estimation from Light FieldsTitus Leistner, Radek Mackowiak, Lynton Ardizzone, Ullrich Köthe et al.CVPR 2022 · 14 citations
- PolarDepth: Monocular Transparent Object Depth from Polar-Physics PriorsWen Dong, Haiyang Mei, Yinglian Ji, Zijun Zhang et al.ICML 2026
- MPI-Flow: Learning Realistic Optical Flow with Multiplane ImagesYingping Liang, Jiaming Liu, Debing Zhang, Ying FuICCV 2023 · 12 citations
- StereOBJ-1M: Large-scale Stereo Image Dataset for 6D Object Pose EstimationXingyu Liu, Shun Iwase, Kris M. KitaniICCV 2021 · 58 citations
- LaRI: Layered Ray Intersections for Single-view 3D Geometric ReasoningRui Li, Biao Zhang, Zhenyu Li, Federico Tombari et al.ICML 2026
