Separating the Wheat from the Chaff: Spatio-Temporal Transformer with View-interweaved Attention for Photon-Efficient Depth Sensing
Letian Yu, Jiaxi Yang, Bo Dong, Qirui Bao, Yuanbo Wang, Felix Heide, Xiaopeng Wei, Xin Yang
Abstract
Time-resolved imaging is an emerging sensing modality that has been shown to enable advanced applications, including remote sensing, fluorescence lifetime imaging, and even non-line-of-sight sensing. Single-photon avalanche diodes (SPADs) outperform relevant time-resolved imaging technologies thanks to their excellent photon sensitivity and superior temporal resolution on the order of tens of picoseconds. The capability of exceeding the sensing limits of conventional cameras for SPADs also draws attention to the photon-efficient imaging area. However, photon-efficient imaging under degraded conditions with low photon counts and low signal-to-background ratio (SBR) still remains an inevitable challenge. In this paper, we propose a spatio-temporal transformer network for photon-efficient imaging under low-flux scenarios. In particular, we introduce a view-interweaved attention mechanism (VIAM) to extract both spatial-view and temporal-view self-attention in each transformer block. We also design an adaptive-weighting scheme to dynamically adjust the weights between different views of self-attention in VIAM for different signal-to-background levels. We extensively validate and demonstrate the effectiveness of our approach on the simulated Middlebury dataset and a specially self-collected dataset with real-world-captured SPAD measurements and well-annotated ground truth depth maps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abfd4008-7d0a-4ee4-a6f5-9428d78fee16Cited by top-tier papers1
Ask how each one uses itBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic SegmentationLi Jiang, Shaoshuai Shi, Zhuotao Tian, Xin Lai et al.ICCV 2021 · 137 citations
- Spatial Pruned Sparse Convolution for Efficient 3D Object DetectionJianhui Liu, Yukang Chen, Xiaoqing Ye, Zhuotao Tian et al.NeurIPS 2022 · 57 citations
- Hybrid Spectral Denoising Transformer with Guided AttentionZeqiang Lai, Chenggang Yan, Ying FuICCV 2023 · 35 citations
Related papers
- Photon-Starved Scene Inference using Single Photon CamerasBhavya Goyal, Mohit GuptaICCV 2021 · 14 citations
- Spiking Vision Transformer with Saccadic AttentionShuai Wang, Malu Zhang, Dehao Zhang, Ammar Belatreche et al.ICLR 2025
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao et al.CVPR 2025
- Spike-RetinexFormer: Rethinking Low-light Image Enhancement with Spiking Neural NetworksHongzhi Wang, Xiubo Liang, Jinxing Han, Weidong GengNeurIPS 2025
- Low-cost SPAD sensing for non-line-of-sight tracking, material classification and depth imagingClara Callenberg, Zheng Shi, Felix Heide, Matthias B. HullinSIGGRAPH 2021 · 55 citations
