Progressive Sparse Local Attention for Video Object Detection
Chaoxu Guo, Bin Fan, Jie Gu, Qian Zhang, Shiming Xiang, Véronique Prinet, Chunhong Pan
Abstract
Transferring image-based object detectors to the domain of videos remains a challenging problem. Previous efforts mostly exploit optical flow to propagate features across frames, aiming to achieve a good trade-off between accuracy and efficiency. However, introducing an extra model to estimate optical flow can significantly increase the overall model size. The gap between optical flow and high-level features can also hinder it from establishing spatial correspondence accurately. Instead of relying on optical flow, this paper proposes a novel module called Progressive Sparse Local Attention (PSLA), which establishes the spatial correspondence between features across frames in a local region with progressively sparser stride and uses the correspondence to propagate features. Based on PSLA, Recursive Feature Updating (RFU) and Dense Feature Transforming (DenseFT) are proposed to model temporal appearance and enrich feature representation respectively in a novel video object detection framework. Experiments on ImageNet VID show that our method achieves the best accuracy compared to existing methods with smaller model size and acceptable runtime speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9040bed-4135-4c85-956b-222311108e35Cited by top-tier papers13
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu et al.AAAI 2021 · 108 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- MAMBA: Multi-level Aggregation via Memory Bank for Video Object DetectionGuanxiong Sun, Yang Hua, Guosheng Hu, Neil RobertsonAAAI 2021 · 73 citations
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao et al.AAAI 2022 · 35 citations
- Distilled Reverse Attention Network for Open-world Compositional Zero-Shot LearningYun Li, Zhe Liu, Saurav Jha, Lina YaoICCV 2023 · 23 citations
Builds on1
Related papers
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi et al.CVPR 2021
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Dual Semantic Fusion Network for Video Object DetectionLijian Lin, Haosheng Chen, Honglun Zhang, Jun Liang et al.ACM MM 2020 · 30 citations
