Exploiting Better Feature Aggregation for Video Object Detection
Liang Han, Pichao Wang, Zhaozheng Yin, Fan Wang, Hao Li
Abstract
Video object detection (VOD) has been a rising topic in recent years due to the challenges such as occlusion, motion blur, etc. To deal with these challenges, feature aggregation from local or global support frames is verified effective. To exploit better feature aggregation, in this paper, we propose two improvements over previous works: a class-constrained spatial-temporal relation network and a correlation-based feature alignment module. For the class constrained spatial-temporal relation network, it operates on object region proposals, and learns two kinds of relations: (1) the dependencies among region proposals of the same object class from support frames sampled in a long time range or even the whole sequence, and (2) spatial relations among proposals of different objects in the target frame. The homogeneity constraint in spatial-temporal relation network not only filters out many defective proposals but also implicitly embeds the traditional post-processing strategies (e.g., Seq-NMS), leading to a unified end-to-end training networks. In the feature alignment module, we propose a correlation based feature alignment method to align the support and target frames for feature aggregation in the temporal domain. Our experiments show that the proposed method improves the accuracy of single-frame detectors significantly, and outperforms previous temporal or spatial relation networks. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset (84.80% with ResNet-101) without any post-processing methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1e918b5a-c5ea-424d-9942-b7f5df2a4788Cited by top-tier papers5
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- Decoupled IoU Regression for Object DetectionYan Gao, Qimeng Wang, Xu Tang, Haochen Wang et al.ACM MM 2021 · 25 citations
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
- TGBFormer: Transformer-GraphFormer Blender Network for Video Object DetectionQiang Qi, Xiao WangAAAI 2025 · 5 citations
- Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object DetectionBingqing Zhang, Sen Wang, Yifan Liu, Brano Kusy et al.ACM MM 2023 · 3 citations
Related papers
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu et al.AAAI 2021 · 108 citations
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong et al.CVPR 2020
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
