Quanta Neural Networks: From Photons to Perception
Varun Sundar, Tianyi Zhang, Sacha Jungerman, Mohit Gupta
Abstract
Quanta image sensors record individual photons, enabling capabilities like imaging in near-complete darkness and ultrahigh-speed videography. Yet, most research on quanta sensors is limited to recovering image intensities. Can we go beyond just imaging, and develop algorithms that can extract high-level scene information from quanta sensors? This could unlock new possibilities in vision systems, offering reliable operation in extreme conditions. The challenge: raw photon streams captured by quanta sensors have fundamentally different characteristics than conventional images, making them incompatible with vision models. One approach is to first transform raw photon streams to conventional-like images, but this is prohibitively expensive in terms of compute, memory, and latency.
We propose quanta neural networks (QNNs) that directly produce downstream task objectives from raw photon streams. Our core proposal is a trainable QNN layer that can seamlessly integrate with existing image-and video-based neural networks, producing quanta counterparts. By avoiding image reconstruction and allocating computational resources on a scene-adaptive basis, QNNs achieve 1-2 orders of magnitude improvements across all efficiency metrics (compute, latency, readout bandwidth) as compared to reconstruction-based quanta vision, while maintaining high task accuracy across a wide gamut of challenging scenarios including low light and rapid motion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b5336ec-f012-4fa2-b1f7-bb1e0cc0a2e4Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- bit2bit: 1-bit quanta video reconstruction via self-supervised photon predictionYehe Liu, Alexander Krull, Hector Basevi, Ales Leonardis et al.NeurIPS 2024 · 4 citations
- Generalized Event CamerasVarun Sundar, Matthew Dutson, Andrei Ardelean, Claudio Bruschini et al.CVPR 2024 · 7 citations
- Quanta burst photographySizhuo Ma, Shantanu Gupta, Arin C. Ulku, Claudio Bruschini et al.SIGGRAPH 2020 · 76 citations
- Eulerian Single-Photon VisionShantanu Gupta, Mohit GuptaICCV 2023 · 6 citations
- gQIR: Generative Quanta Image ReconstructionAryan Garg, Sizhuo Ma, Mohit GuptaCVPR 2026
