Efficient Depth Estimation for Unstable Stereo Camera Systems on AR Glasses
Yongfan Liu, Hyoukjun Kwon
摘要
Stereo depth estimation is a fundamental component in augmented reality (AR), which requires low latency for realtime processing. However, preprocessing such as rectification and non-ML computations such as cost volume require significant amount of latency exceeding that of an ML model itself, which hinders the real-time processing required by AR. Therefore, we develop alternative approaches to the rectification and cost volume that consider ML acceleration (GPU and NPUs) in recent hardware. For pre-processing, we eliminate it by introducing homography matrix prediction network with a rectification positional encoding (RPE), which delivers both low latency and robustness to unrectified images. For cost volume, we replace it with a grouppointwise convolution-based operator and approximation of cosine similarity based on layernorm and dot product. Based on our approaches, we develop MultiHeadDepth (replacing cost volume) and HomoDepth (MultiHeadDepth + removing pre-processing) models. MultiHeadDepth provides 11.8-30.3% improvements in accuracy and 22.9-25.2% reduction in latency compared to a state-of-the-art depth estimation model for AR glasses from industry. Ho-moDepth, which can directly process unrectified images, reduces the end-to-end latency by 44.5%. We also introduce a multi-task learning method to handle misaligned stereo inputs on HomoDepth, which reduces the AbsRel error by 10.0-24.3%. The overall results demonstrate the efficacy of our approaches, which not only reduce the inference latency but also improve the model performance. Our code is available at https://github .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine PerceptionXiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters 等ICCV 2023 · 被引用 145 次
- One shot 3D photographyJohannes Kopf, Kevin Matzen, Suhib Alsisan, Ocean Quigley 等SIGGRAPH 2020 · 被引用 65 次
- Selective-Stereo: Adaptive Frequency Information Selection for Stereo MatchingXianqi Wang, Gangwei Xu, Hao Jia, Xin YangCVPR 2024 · 被引用 64 次
- MobiDepth: real-time depth estimation using on-device dual camerasJinrui Zhang, Huan Yang, Ju Ren, Deyu Zhang 等MobiCom 2022 · 被引用 24 次
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova 等CVPR 2023
相关 Paper
- A Practical Stereo Depth System for Smart GlassesJialiang Wang, Daniel Scharstein, Akash Bapat, Kevin Blackburn-Matzen 等CVPR 2023
- Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationRui Gong, Kim-Hui Yap, Weide Liu, Xulei Yang 等CVPR 2025
- Integrating Both Parallax and Latency Compensation into Video See-through Head-mounted DisplayAtsushi Ishihara, Hiroyuki Aga, Yasuko Ishihara, Hirotake Ichikawa 等IEEE VR 2023 · 被引用 16 次
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 被引用 71 次
- HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented RealityShulin Zhao, Haibo Zhang, Cyan Subhra Mishra, Sandeepa Bhuyan 等MICRO 2021 · 被引用 22 次
