Sequence Level Semantics Aggregation for Video Object Detection
Haiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang
摘要
Video objection detection (VID) has been a rising research direction in recent years. A central issue of VID is the appearance degradation of video frames caused by fast motion. This problem is essentially ill-posed for a single frame. Therefore, aggregating features from other frames becomes a natural choice. Existing methods rely heavily on optical flow or recurrent neural networks for feature aggregation. However, these methods emphasize more on the temporally nearby frames. In this work, we argue that aggregating features in the full-sequence level will lead to more discriminative and robust features for video object detection. To achieve this goal, we devise a novel Sequence Level Semantics Aggregation (SELSA) module. We further demonstrate the close relationship between the proposed method and the classic spectral clustering method, providing a novel view for understanding the VID problem. We test the proposed method on the ImageNet VID and the EPIC KITCHENS dataset and achieve new state-of-the-art results. Our method does not need complicated postprocessing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 被引用 180 次
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 被引用 171 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li 等ICCV 2021 · 被引用 124 次
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu 等AAAI 2021 · 被引用 108 次
相关 Paper
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang 等ACM MM 2020 · 被引用 37 次
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- InstMove: Instance Motion for Object-centric Video SegmentationQihao Liu, Junfeng Wu, Yi Jiang, Xiang Bai 等CVPR 2023
- CompFeat: Comprehensive Feature Aggregation for Video Instance SegmentationYang Fu, Linjie Yang, Ding Liu, Thomas S. Huang 等AAAI 2021 · 被引用 77 次
