DepthFocus: Controllable Depth Estimation for See-Through Scenes
junhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min, YOUNGPIL JEON, Minyong Choi
摘要
Depth in the real world is rarely singular. Transmissive materials create layered ambiguities that confound conventional perception systems. Existing models remain passive, attempting to estimate static depth maps anchored to the nearest surface, while humans actively shift focus to perceive a desired depth. We introduce DepthFocus, a steerable Vision Transformer that redefines stereo depth estimation as intent-driven control. Conditioned on a scalar depth preference, the model dynamically adapts its computation to focus on the intended depth, enabling selective perception within complex scenes. The training primarily leverages our newly constructed 500k multi-layered synthetic dataset, designed to capture diverse see-through effects. DepthFocus not only achieves state-of-the-art performance on conventional single-depth benchmarks like BOOSTER, a dataset notably rich in transparent and reflective objects, but also quantitatively demonstrates intent-aligned estimation on our newly proposed real and synthetic multi-depth datasets. Moreover, it exhibits strong generalization capabilities on unseen see-through scenes, underscoring its robustness as a significant step toward active and human-like 3D perception.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
相关 Paper
- Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth EstimationHongyu Wen, Yiming Zuo, Venkat Subramanian, Patrick Chen 等ICCV 2025 · 被引用 1 次
- DualFocus: Depth from Focus with Spatio-Focal Dual Variational ConstraintsSungmin Woo, Sangyoun LeeNeurIPS 2025 · 被引用 2 次
- Learning Depth Estimation for Transparent and Mirror SurfacesAlex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi 等ICCV 2023 · 被引用 41 次
- Video See-Through Mixed Reality with Focus CuesChristoph Ebner, Shohei Mori, Peter Mohr, Yifan Peng 等IEEE VR 2022 · 被引用 24 次
- SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined GroupingHongyu Wen, Jia DengCVPR 2026
