240FPS Stereo Vision from Monocular Mixed Spikes
Yeliduosi Xiaokaiti, Yakun Chang, Yang Bai, Zhaojun Huang, Peiqi Duan, Boxin Shi
Abstract
Stereo vision is fundamental for enabling machines to perceive and interact with the world. While monocular stereo methods offer hardware compactness, they struggle with generalization due to reliance on data-driven priors. Binocular and multi-view systems improve accuracy but incur higher hardware complexity and data inefficiency. In this paper, we introduce a monocular solution for high-framerate stereo vision via temporal optical modulation. The modulation directs light from two views onto a single sensor in a mixed manner, while periodically attenuating one view at 60 Hz. To capture the temporal variations introduced by this modulation, we employ a high-speed spike camera that records the mixed scene as temporally dense spikes. The high temporal resolution of these spikes enables the construction of a linear system for efficient binocular video decoupling. Consequently, we introduce a two-stage decoding methodology for achieving high-quality stereo vision: An efficient least-squares-based baseline reconstruction followed by a deep learning refinement module. Experimental results demonstrate that our approach achieves 240FPS binocular video reconstruction with superior accuracy compared to monocular systems, while maintaining the hardware compactness and data efficiency. Code is available at https://github.com/yongqiye00/MonoSpikeStereo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b6dff9c-8886-473e-946f-7c4f0daf90b1Builds on12
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai et al.CVPR 2022 · 294 citations
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 16 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
- Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking CamerasBin Fan, Jiaoyang Yin, Yuchao Dai, Chao Xu et al.NeurIPS 2024 · 7 citations
Related papers
- Enhancing Motion Deblurring in High-Speed Scenes with Spike StreamsShiyan Chen, Jiyuan Zhang, Yajing Zheng, Tiejun Huang et al.NeurIPS 2023 · 21 citations
- BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo VisionYiqian Chang, Haoran Xu, Qinghong Ye, Jianing Li et al.AAAI 2026
- SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike StreamsZhuoheng Gao, Yihao Li, Jiyao Zhang, Rui Zhao et al.ICLR 2026 · 2 citations
- Optical Flow Estimation for Spiking CameraLiwen Hu, Rui Zhao, Ziluo Ding, Lei Ma et al.CVPR 2022 · 48 citations
- HFR and HDR Video from Multi-Attenuated Spikes Using a Rapidly Rotating SpokeND FilterYakun Chang, Zhaojun Huang, Siqi Yang, Yeliduosi Xiaokaiti et al.CVPR 2026
