Quanta Neural Networks: From Photons to Perception
Varun Sundar, Tianyi Zhang, Sacha Jungerman, Mohit Gupta
摘要
Quanta image sensors record individual photons, enabling capabilities like imaging in near-complete darkness and ultrahigh-speed videography. Yet, most research on quanta sensors is limited to recovering image intensities. Can we go beyond just imaging, and develop algorithms that can extract high-level scene information from quanta sensors? This could unlock new possibilities in vision systems, offering reliable operation in extreme conditions. The challenge: raw photon streams captured by quanta sensors have fundamentally different characteristics than conventional images, making them incompatible with vision models. One approach is to first transform raw photon streams to conventional-like images, but this is prohibitively expensive in terms of compute, memory, and latency.
We propose quanta neural networks (QNNs) that directly produce downstream task objectives from raw photon streams. Our core proposal is a trainable QNN layer that can seamlessly integrate with existing image-and video-based neural networks, producing quanta counterparts. By avoiding image reconstruction and allocating computational resources on a scene-adaptive basis, QNNs achieve 1-2 orders of magnitude improvements across all efficiency metrics (compute, latency, readout bandwidth) as compared to reconstruction-based quanta vision, while maintaining high task accuracy across a wide gamut of challenging scenarios including low light and rapid motion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- bit2bit: 1-bit quanta video reconstruction via self-supervised photon predictionYehe Liu, Alexander Krull, Hector Basevi, Ales Leonardis 等NeurIPS 2024 · 被引用 4 次
- Generalized Event CamerasVarun Sundar, Matthew Dutson, Andrei Ardelean, Claudio Bruschini 等CVPR 2024 · 被引用 7 次
- Quanta burst photographySizhuo Ma, Shantanu Gupta, Arin C. Ulku, Claudio Bruschini 等SIGGRAPH 2020 · 被引用 76 次
- Eulerian Single-Photon VisionShantanu Gupta, Mohit GuptaICCV 2023 · 被引用 6 次
- gQIR: Generative Quanta Image ReconstructionAryan Garg, Sizhuo Ma, Mohit GuptaCVPR 2026
