Context Enhanced Transformer for Single Image Object Detection in Video Data
Seungjun An, Seonghoon Park, Gyeongnyeon Kim, Jeongyeol Baek, Byeongwon Lee, Seungryong Kim
Abstract
With the increasing importance of video data in real-world applications, there is a rising need for efficient object detection methods that utilize temporal information. While existing video object detection (VOD) techniques employ various strategies to address this challenge, they typically depend on locally adjacent frames or randomly sampled images within a clip. Although recent Transformer-based VOD methods have shown promising results, their reliance on multiple inputs and additional network complexity to incorporate temporal information limits their practical applicability. In this paper, we propose a novel approach to single image object detection, called Context Enhanced TRansformer (CETR), by incorporating temporal context into DETR using a newly designed memory module. To efficiently store temporal information, we construct a class-wise memory that collects contextual information across data. Additionally, we present a classification-based sampling technique to selectively utilize the relevant memory for the current image. In the testing, We introduce a test-time memory adaptation method that updates individual memory functions by considering the test distribution. Experiments with CityCam and ImageNet VID datasets exhibit the efficiency of the framework on various video systems. The project page and code will be made available at: https://ku-cvlab.github.io/CETR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- TGBFormer: Transformer-GraphFormer Blender Network for Video Object DetectionQiang Qi, Xiao WangAAAI 2025 · 5 citations
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- MSTDiff: Multiscale-Aware Transformer Diffusion Network for Video Object DetectionQiang Qi, Wenqi Shang, Xiao Wang, Yanjie Liang et al.AAAI 2026
Builds on19
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram et al.AAAI 2024 · 4 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
- InstanceFormer: An Online Video Instance Segmentation FrameworkRajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh et al.AAAI 2023 · 16 citations
