ActiveZero: Mixed Domain Learning for Active Stereovision with Zero Annotation
Isabella Liu, Edward Yang, Jianyu Tao, Rui Chen, Xiaoshuai Zhang, Qing Ran, Zhu Liu, Hao Su
Abstract
Traditional depth sensors generate accurate real world depth estimates that surpass even the most advanced learning approaches trained only on simulation domains. Since ground truth depth is readily available in the simulation domain but quite difficult to obtain in the real domain, we propose a method that leverages the best of both worlds. In this work we present a new framework, ActiveZero, which is a mixed domain learning solution for active stereovision systems that requires no real world depth annotation. First, we demonstrate the transferability of our method to out-of-distribution real data by using a mixed domain learning strategy. In the simulation domain, we use a combination of supervised disparity loss and self-supervised losses on a shape primitives dataset. By contrast, in the real domain, we only use self-supervised losses on a dataset that is out-of-distribution from either training simulation data or test real data. Second, our method introduces a novel self-supervised loss called temporal IR reprojection to increase the robustness and accuracy of our reprojections in hard-to-perceive regions. Finally, we show how the method can be trained end-to-end and that each module is important for attaining the end result. Extensive qualitative and quantitative evaluations on real data demonstrate state of the art results that can even beat a commercial depth sensor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Active Stereo Without Pattern ProjectorLuca Bartolomei, Matteo Poggi, Fabio Tosi, Andrea Conti et al.ICCV 2023 · 12 citations
- GS-ASM: 2DGS-Supervised Active Stereo MatchingZhengling Wu, Rongfeng Lu, Quan Chen, Longjian Zeng et al.CVPR 2026
Builds on3
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai et al.CVPR 2020
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo MatchingRui Liu, Chengxi Yang, Wenxiu Sun, Xiaogang Wang et al.CVPR 2020
Related papers
- Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsWeilong Yan, Ming Li, Haipeng Li, Shuwei Shao et al.CVPR 2025
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth EstimationMohammad Mahdi Johari, Camilla Carta, François FleuretICCV 2021 · 12 citations
- Self-supervised Multi-view Stereo via Inter and Intra Network Pseudo DepthKe Qiu, Yawen Lai, Shiyi Liu, Ronggang WangACM MM 2022 · 9 citations
- Robust Geometry-Preserving Depth Estimation Using Differentiable RenderingChi Zhang, Wei Yin, Gang Yu, Zhibin Wang et al.ICCV 2023 · 7 citations
