SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
Gaole Dai, Menghang Dong, Rongyu Zhang, Ruichuan An, Tiejun Huang, Shanghang Zhang
Abstract
The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the functions of cone and rod cells: cones are primarily responsible for color perception, while rods are specialized in detecting motion, particularly variations in light intensity. These two distinct modalities of visual information are integrated and processed within the visual cortex, thereby enhancing the robustness of the human visual system. Inspired by this biological mechanism, modern hardware systems have evolved to include not only color-sensitive RGB cameras but also motion-sensitive Dynamic Visual Systems, such as spike cameras. Building upon these advancements, this study seeks to emulate the human visual system by integrating decomposed multi-modal visual inputs with modern latent-space generative frameworks. We named it SpikeGen. We evaluate its performance across various spike-RGB tasks, including conditional image and video deblurring, dense frame reconstruction from spike streams, and high-speed scene novel-view synthesis. Supported by extensive experiments, we demonstrate that leveraging the latent space manipulation capabilities of generative models enables an effective synergistic enhancement of different visual modalities, addressing spatial sparsity in spike inputs and temporal sparsity in RGB inputs. Codes and pretrained weights are avaliable in https://github.com/zhenwuweihe/SpikeGen.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77ae3102-f2ec-4352-b199-7444e9f44fa2Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Enhancing Motion Deblurring in High-Speed Scenes with Spike StreamsShiyan Chen, Jiyuan Zhang, Yajing Zheng, Tiejun Huang et al.NeurIPS 2023 · 21 citations
- Retina-Like Visual Image Reconstruction via Spiking Neural ModelLin Zhu, Siwei Dong, Jianing Li, Tiejun Huang et al.CVPR 2020
- SVFI: Spiking-Based Video Frame Interpolation for High-Speed MotionLujie Xia, Jing Zhao, Ruiqin Xiong, Tiejun HuangAAAI 2023 · 14 citations
- 240FPS Stereo Vision from Monocular Mixed SpikesYeliduosi Xiaokaiti, Yakun Chang, Yang Bai, Zhaojun Huang et al.CVPR 2026
- SpikeGS: 3D Gaussian Splatting from Spike Streams with High-Speed Camera MotionJiyuan Zhang, Kang Chen, Shiyan Chen, Yajing Zheng et al.ACM MM 2024 · 8 citations
