Efficient Depth Estimation for Unstable Stereo Camera Systems on AR Glasses
Yongfan Liu, Hyoukjun Kwon
Abstract
Stereo depth estimation is a fundamental component in augmented reality (AR), which requires low latency for realtime processing. However, preprocessing such as rectification and non-ML computations such as cost volume require significant amount of latency exceeding that of an ML model itself, which hinders the real-time processing required by AR. Therefore, we develop alternative approaches to the rectification and cost volume that consider ML acceleration (GPU and NPUs) in recent hardware. For pre-processing, we eliminate it by introducing homography matrix prediction network with a rectification positional encoding (RPE), which delivers both low latency and robustness to unrectified images. For cost volume, we replace it with a grouppointwise convolution-based operator and approximation of cosine similarity based on layernorm and dot product. Based on our approaches, we develop MultiHeadDepth (replacing cost volume) and HomoDepth (MultiHeadDepth + removing pre-processing) models. MultiHeadDepth provides 11.8-30.3% improvements in accuracy and 22.9-25.2% reduction in latency compared to a state-of-the-art depth estimation model for AR glasses from industry. Ho-moDepth, which can directly process unrectified images, reduces the end-to-end latency by 44.5%. We also introduce a multi-task learning method to handle misaligned stereo inputs on HomoDepth, which reduces the AbsRel error by 10.0-24.3%. The overall results demonstrate the efficacy of our approaches, which not only reduce the inference latency but also improve the model performance. Our code is available at https://github .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine PerceptionXiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters et al.ICCV 2023 · 145 citations
- One shot 3D photographyJohannes Kopf, Kevin Matzen, Suhib Alsisan, Ocean Quigley et al.SIGGRAPH 2020 · 65 citations
- Selective-Stereo: Adaptive Frequency Information Selection for Stereo MatchingXianqi Wang, Gangwei Xu, Hao Jia, Xin YangCVPR 2024 · 64 citations
- MobiDepth: real-time depth estimation using on-device dual camerasJinrui Zhang, Huan Yang, Ju Ren, Deyu Zhang et al.MobiCom 2022 · 24 citations
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova et al.CVPR 2023
Related papers
- A Practical Stereo Depth System for Smart GlassesJialiang Wang, Daniel Scharstein, Akash Bapat, Kevin Blackburn-Matzen et al.CVPR 2023
- Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationRui Gong, Kim-Hui Yap, Weide Liu, Xulei Yang et al.CVPR 2025
- Integrating Both Parallax and Latency Compensation into Video See-through Head-mounted DisplayAtsushi Ishihara, Hiroyuki Aga, Yasuko Ishihara, Hirotake Ichikawa et al.IEEE VR 2023 · 16 citations
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 71 citations
- HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented RealityShulin Zhao, Haibo Zhang, Cyan Subhra Mishra, Sandeepa Bhuyan et al.MICRO 2021 · 22 citations
