Temporal Context Enhanced Feature Aggregation for Video Object Detection
Fei He, Naiyu Gao, Qiaozhe Li, Senyao Du, Xin Zhao, Kaiqi Huang
Abstract
Video object detection is a challenging task because of the presence of appearance deterioration in certain video frames. One typical solution is to aggregate neighboring features to enhance per-frame appearance features. However, such a method ignores the temporal relations between the aggregated frames, which is critical for improving video recognition accuracy. To handle the appearance deterioration problem, this paper proposes a temporal context enhanced network (TCENet) to exploit temporal context information by temporal aggregation for video object detection. To handle the displacement of the objects in videos, a novel DeformAlign module is proposed to align the spatial features from frame to frame. Instead of adopting a fixed-length window fusion strategy, a temporal stride predictor is proposed to adaptively select video frames for aggregation, which facilitates exploiting variable temporal information and requiring fewer video frames for aggregation to achieve better results. Our TCENet achieves state-of-the-art performance on the Ima-geNet VID dataset and has a faster runtime. Without bellsand-whistles, our TCENet achieves 80.3% mAP by only aggregating 3 frames.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao et al.AAAI 2022 · 35 citations
- Query-Memory Re-Aggregation for Weakly-supervised Video Object SegmentationFanchao Lin, Hongtao Xie, Yan Li, Yongdong ZhangAAAI 2021 · 26 citations
- InsPro: Propagating Instance Query and Proposal for Online Video Instance SegmentationFei He, Haoyang Zhang, Naiyu Gao, Jian Jia et al.NeurIPS 2022 · 23 citations
- TempNet: Online Semantic Segmentation on Large-scale Point Cloud SeriesYunsong Zhou, Hongzi Zhu, Chunqin Li, Tiankai Cui et al.ICCV 2021 · 6 citations
Builds on2
Related papers
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 171 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu et al.AAAI 2021 · 108 citations
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
