DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging
Xingjian Jiang, Lishun Wang, Ping Wang, Xin Yuan
Abstract
Video snapshot compressive imaging (SCI) offers a promising alternative to high-speed cameras by encoding multiple frames into a single 2D measurement. However, SCI requires algorithms to reconstruct the high-speed video, and as resolution increases, reconstruction becomes computationally expensive and memory-intensive. Much of the resource is wasted on recovering large background regions that contain little useful information, highlighting the need for selective, object-driven reconstruction. Existing object detectors struggle to perform accurately on SCI measurements due to the spatial-temporal aliasing introduced by coded exposure. To address this challenge, we propose De-tectSCI, the first framework enabling object-guided regionof-interest (ROI) reconstruction for high-resolution SCI. The inside detector comprises two key components: an encoder built from weight-sharing Mamba-Implicit Modules (MIM) for progressive feature refinement, and a Frequency Mamba (FM) module dedicated to frequency-aware query selection. MIM enhances features via multi-scale dilated convolutions and implicit representations, while FM restores discriminative details by decomposing and reweighting frequency bands. Experiments on the SportsMOT dataset show that DetectSCI achieves 80.9 Average Precision (AP), surpassing the best CNN-based detector by at least 2.8 AP and the best Transformer-based detector by at least 4.1 AP, while maintaining comparable efficiency. Code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa9b5e1d-570b-4d2d-8cc4-197e98eb3aaaBuilds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- YOLOv10: Real-Time End-to-End Object DetectionAo Wang, Hui Chen, Lihao Liu, Kai Chen et al.NeurIPS 2024 · 6,113 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 2,652 citations
Related papers
- MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive SensingZhengjue Wang, Hao Zhang, Ziheng Cheng, Bo Chen et al.CVPR 2021
- EfficientSCI: Densely Connected Network with Space-time Factorization for Large-scale Video Snapshot Compressive ImagingLishun Wang, Miao Cao, Xin YuanCVPR 2023
- MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive ImagingZhenghao Pan, Haijin Zeng, Jiezhang Cao, Yongyong Chen et al.NeurIPS 2024 · 12 citations
- Towards Real-time Video Compressive Sensing on Mobile DevicesMiao Cao, Lishun Wang, Huan Wang, Guoqing Wang et al.ACM MM 2024 · 4 citations
- Memory-Efficient Network for Large-Scale Video Compressive SensingZiheng Cheng, Bo Chen, Guanliang Liu, Hao Zhang et al.CVPR 2021
