End-to-End Video Object Detection with Spatial-Temporal Transformers
Lu He, Qianyu Zhou, Xiangtai Li, Li Niu, Guangliang Cheng, Xiao Li, Wenxuan Liu, Yunhai Tong, Lizhuang Ma, Liqing Zhang
摘要
Recently, DETR and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors. However, their performance on Video Object Detection (VOD) has not been well explored. In this paper, we present TransVOD, an end-to-end video object detection model based on a spatial-temporal Transformer architecture. The goal of this paper is to streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks. Besides, benefited from the object query design in DETR, our method does not need complicated post-processing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. In particular, we present temporal Transformer to aggregate both the spatial object queries and the feature memories of each frame. Our temporal Transformer consists of three components: Temporal Deformable Transformer Encoder (TDTE) to encode the multiple frame spatial details, Temporal Query Encoder (TQE) to fuse object queries, and Temporal Deformable Transformer Decoder to obtain current frame detection results. These designs boost the strong baseline deformable DETR by a significant margin (3%-4% mAP) on the ImageNet VID dataset. TransVOD yields comparable results performance on the benchmark of ImageNet VID. We hope our TransVOD can provide a new perspective for video object detection. Code will be made publicly available at https://github.com/SJTU-LuHe/TransVOD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等CVPR 2022 · 被引用 87 次
- YOLOV: Making Still Image Object Detectors Great at Video Object DetectionYuheng Shi, Naiyan Wang, Xiaojie GuoAAAI 2023 · 被引用 83 次
- READ: Large-Scale Neural Scene Rendering for Autonomous DrivingZhuopeng Li, Lu Li, Jianke ZhuAAAI 2023 · 被引用 78 次
- Edge-Assisted On-Device Model Update for Video Analytics in Adverse EnvironmentsYuxin Kong, Peng Yang, Yan ChengACM MM 2023 · 被引用 40 次
- Explore Spatio-temporal Aggregation for Insubstantial Object Detection: Benchmark Dataset and BaselineKailai Zhou, Yibo Wang, Tao Lv, Yunqian Li 等CVPR 2022 · 被引用 21 次
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
相关 Paper
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object DetectionQiang Qi, Wenqi Shang, Xiao Wang, Yanjie Liang 等AAAI 2026
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram 等AAAI 2024 · 被引用 4 次
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du 等AAAI 2020 · 被引用 40 次
