Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cues
Yuhui Liu, Liangxun Ou, Qiang Fu, Hadi Amata, Wolfgang Heidrich, Yifan Peng
Abstract
Extracting high-fidelity RGBD information from twodimensional (2D) images is essential for various visual computing applications. Stereo imaging, as a reliable passive imaging technique for obtaining three-dimensional (3D) scene information, has benefited greatly from deep learning advancements. However, existing stereo depth estimation algorithms struggle to perceive high-frequency information and resolve high-resolution depth maps in realistic camera settings with large depth variations. These algorithms commonly neglect the hardware parameter configuration, limiting the potential for achieving optimal solutions solely through software-based design strategies.
This work presents a hardware-software co-designed RGBD imaging framework that leverages both stereo and focus cues to reconstruct texture-rich color images along with detailed depth maps over a wide depth range. A pair of rank-2 parameterized diffractive optical elements (DOEs) is employed to encode perpendicular complementary information optically during stereo acquisitions. Additionally, we employ an IGEV-UNet-fused neural network tailored to the proposed rank-2 encoding for stereo matching and image reconstruction. Through prototyping a stereo camera with customized DOEs, our deep stereo imaging paradigm has demonstrated superior performance over existing monocular and stereo imaging systems in both image PSNR by 2.96 dB gain and depth accuracy in highfrequency details across distances from 0.67 to 8 meters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c49dee59-9985-49c3-a83b-5a8946660c3eBuilds on11
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- Attention Concatenation Volume for Accurate and Efficient Stereo MatchingGangwei Xu, Junda Cheng, Peng Guo, Xin YangCVPR 2022 · 265 citations
- Single-shot Hyperspectral-Depth Imaging with Learned Diffractive OpticsSeung-Hwan Baek, Hayato Ikoma, Daniel S. Jeon, Yuqi Li et al.ICCV 2021 · 109 citations
- Selective-Stereo: Adaptive Frequency Information Selection for Stereo MatchingXianqi Wang, Gangwei Xu, Hao Jia, Xin YangCVPR 2024 · 64 citations
- Seeing through obstructions with diffractive cloakingZheng Shi, Yuval Bahat, Seung-Hwan Baek, Qiang Fu et al.SIGGRAPH 2022 · 34 citations
Related papers
- CodedStereo: Learned Phase Masks for Large Depth-of-Field StereoShiyu Tan, Yicheng Wu, Shoou-I Yu, Ashok VeeraraghavanCVPR 2021
- DPS-Net: Deep Polarimetric Stereo Depth EstimationChaoran Tian, Weihong Pan, Zimo Wang, Mao Mao et al.ICCV 2023 · 24 citations
- MVS2D: Efficient Multiview Stereo via Attention-Driven 2D ConvolutionsZhenpei Yang, Zhile Ren, Qi Shan, Qixing HuangCVPR 2022 · 43 citations
- Polka Lines: Learning Structured Illumination and Reconstruction for Active StereoSeung-Hwan Baek, Felix HeideCVPR 2021
- 240FPS Stereo Vision from Monocular Mixed SpikesYeliduosi Xiaokaiti, Yakun Chang, Yang Bai, Zhaojun Huang et al.CVPR 2026
