DepthFocus: Controllable Depth Estimation for See-Through Scenes
junhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min, YOUNGPIL JEON, Minyong Choi
Abstract
Depth in the real world is rarely singular. Transmissive materials create layered ambiguities that confound conventional perception systems. Existing models remain passive, attempting to estimate static depth maps anchored to the nearest surface, while humans actively shift focus to perceive a desired depth. We introduce DepthFocus, a steerable Vision Transformer that redefines stereo depth estimation as intent-driven control. Conditioned on a scalar depth preference, the model dynamically adapts its computation to focus on the intended depth, enabling selective perception within complex scenes. The training primarily leverages our newly constructed 500k multi-layered synthetic dataset, designed to capture diverse see-through effects. DepthFocus not only achieves state-of-the-art performance on conventional single-depth benchmarks like BOOSTER, a dataset notably rich in transparent and reflective objects, but also quantitatively demonstrates intent-aligned estimation on our newly proposed real and synthetic multi-depth datasets. Moreover, it exhibits strong generalization capabilities on unseen see-through scenes, underscoring its robustness as a significant step toward active and human-like 3D perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c37dce23-ab25-468a-8f60-d7c75ae47703Builds on35
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
Related papers
- Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth EstimationHongyu Wen, Yiming Zuo, Venkat Subramanian, Patrick Chen et al.ICCV 2025 · 1 citation
- DualFocus: Depth from Focus with Spatio-Focal Dual Variational ConstraintsSungmin Woo, Sangyoun LeeNeurIPS 2025 · 2 citations
- Learning Depth Estimation for Transparent and Mirror SurfacesAlex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi et al.ICCV 2023 · 41 citations
- Video See-Through Mixed Reality with Focus CuesChristoph Ebner, Shohei Mori, Peter Mohr, Yifan Peng et al.IEEE VR 2022 · 24 citations
- SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined GroupingHongyu Wen, Jia DengCVPR 2026
