MobiDepth: real-time depth estimation using on-device dual cameras
Jinrui Zhang, Huan Yang, Ju Ren, Deyu Zhang, Bangwen He, Ting Cao, Yuanchun Li, Yaoxue Zhang, Yunxin Liu
Abstract
Real-time depth estimation is critical for the increasingly popular augmented reality and virtual reality applications on mobile devices. Yet existing solutions are insufficient as they require expensive depth sensors or motion of the device, or have a high latency. We propose MobiDepth, a real-time depth estimation system using the widely-available on-device dual cameras. While binocular depth estimation is a mature technique, it is challenging to realize the technique on commodity mobile devices due to the different focal lengths and unsynchronized frame flows of the on-device dual cameras and the heavy stereo-matching algorithm.
To address the challenges, MobiDepth integrates three novel techniques: 1) iterative field-of-view cropping, which crops the field-of-views of the dual cameras to achieve the equivalent focal lengths for accurate epipolar rectification; 2) heterogeneous camera synchronization, which synchronizes the frame flows captured by the dual cameras to avoid the displacement of moving objects across the frames in the same pair; 3) mobile GPU-friendly stereo matching, which effectively reduces the latency of stereo matching on a mobile GPU. We implement MobiDepth on multiple commodity mobile devices and conduct comprehensive evaluations. Experimental results show that MobiDepth achieves real-time depth estimation of 22 frames per second with a significantly reduced depth-estimation error compared with the baselines. Using MobiDepth, we further build an example application of 3D pose estimation, which significantly outperforms the state-of-the-art 3D pose-estimation method, reducing the pose-estimation latency and error by up to 57.1% and 29.5%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 480fab73-8aaa-4512-8781-83f9c27e9bf6Cited by top-tier papers2
- Asymmetric Dual-Lens Video DeblurringZeyu Xiao, Xinchao WangNeurIPS 2025 · 2 citations
- Efficient Depth Estimation for Unstable Stereo Camera Systems on AR GlassesYongfan Liu, Hyoukjun KwonCVPR 2025
Builds on5
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 71 citations
- Romou: rapidly generate high-performance tensor kernels for mobile GPUsRendong Liang, Ting Cao, Jicheng Wen, Manni Wang et al.MobiCom 2022 · 13 citations
- HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo MatchingVladimir Tankovich, Christian Hane, Yinda Zhang, Adarsh Kowdle et al.CVPR 2021
- Lite-HRNet: A Lightweight High-Resolution NetworkChangqian Yu, Bin Xiao, Changxin Gao, Lu Yuan et al.CVPR 2021
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi et al.CVPR 2020
Related papers
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah et al.ICCV 2025 · 1 citation
- EasyREG: Easy Depth-Based Markerless Registration and Tracking using Augmented Reality Device for Surgical GuidanceYue Yang, Christoph Leuze, Brian A. Hargreaves, Bruce Daniel et al.IEEE VR 2026 · 4 citations
- Lite Pose: Efficient Architecture Design for 2D Human Pose EstimationYihan Wang, Muyang Li, Han Cai, Wei-Ming Chen et al.CVPR 2022 · 117 citations
- Learning Single Camera Depth Estimation Using Dual-PixelsRahul Garg, Neal Wadhwa, Sameer Ansari, Jonathan T. BarronICCV 2019 · 123 citations
- RePoseD: Efficient Relative Pose Estimation With Known Depth InformationYaqing Ding, Viktor Kocur, Václav Vávra, Zuzana Berger Haladová et al.ICCV 2025 · 2 citations
