Relation Distillation Networks for Video Object Detection
Jiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou, Houqiang Li, Tao Mei
Abstract
It has been well recognized that modeling object-to-object relations would be helpful for object detection. Nevertheless, the problem is not trivial especially when exploring the interactions between objects to boost video object detectors. The difficulty originates from the aspect that reliable object relations in a video should depend on not only the objects in the present frame but also all the supportive objects extracted over a long range span of the video. In this paper, we introduce a new design to capture the interactions across the objects in spatio-temporal context. Specifically, we present Relation Distillation Networks (RDN) --- a new architecture that novelly aggregates and propagates object relation to augment object features for detection. Technically, object proposals are first generated via Region Proposal Networks (RPN). RDN then, on one hand, models object relation via multi-stage reasoning, and on the other, progressively distills relation through refining supportive object proposals with high objectness scores in a cascaded manner. The learnt relation verifies the efficacy on both improving object detection in each frame and box linking across frames. Extensive experiments are conducted on ImageNet VID dataset, and superior results are reported when comparing to state-of-the-art methods. More remarkably, our RDN achieves 81.8% and 83.2% mAP with ResNet-101 and ResNeXt-101, respectively. When further equipped with linking and rescoring, we obtain to-date the best reported mAP of 83.8% and 84.7%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 734c005f-507d-4a77-a31f-53547b9c326cCited by top-tier papers38
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 171 citations
- Online Knowledge Distillation for Efficient Pose EstimationZheng Li, Jingwen Ye, Mingli Song, Ying Huang et al.ICCV 2021 · 123 citations
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu et al.AAAI 2021 · 108 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
Related papers
- Object Dynamics Distillation for Scene Decomposition and RepresentationQu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangICLR 2022 · 7 citations
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- Dual Semantic Fusion Network for Video Object DetectionLijian Lin, Haosheng Chen, Honglun Zhang, Jun Liang et al.ACM MM 2020 · 30 citations
- LGD: Label-Guided Self-Distillation for Object DetectionPeizhen Zhang, Zijian Kang, Tong Yang, Xiangyu Zhang et al.AAAI 2022 · 38 citations
