Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection
Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, Jonathan Huang
Abstract
In static monitoring cameras, useful contextual information can stretch far beyond the few seconds typical video understanding models might see: subjects may exhibit similar behavior over multiple days, and background objects remain static. Due to power and storage constraints, sampling frequencies are low, often no faster than one frame per second, and sometimes are irregular due to the use of a motion trigger. In order to perform well in this setting, models must be robust to irregular sampling rates. In this paper we propose a method that leverages temporal context from the unlabeled frames of a novel camera to improve performance at that camera. Specifically, we propose an attention-based approach that allows our model, Context R-CNN, to index into a long term memory bank constructed on a per-camera basis and aggregate contextual features from other frames to boost object detection performance on the current frame. We apply Context R-CNN to two settings: (1) species detection using camera traps, and (2) vehicle detection in traffic cameras, showing in both settings that Context R-CNN leads to performance gains over strong baselines. Moreover, we show that increasing the contextual time horizon leads to improved results. When applied to camera trap data from the Snapshot Serengeti dataset, Context R-CNN with context from up to a month of images outperforms a single-frame baseline by 17.9% mAP, and outperforms S3D (a 3d convolution based baseline) by 11.2% mAP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f8670a2-e901-4e27-bb02-6d2fa9c71053Cited by top-tier papers13
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu et al.MobiCom 2021 · 103 citations
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong et al.CVPR 2022 · 83 citations
- Exploiting Temporal Relations on Radar Perception for Autonomous DrivingPeizhao Li, Pu Wang, Karl Berntorp, Hongfu LiuCVPR 2022 · 50 citations
Builds on6
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
Related papers
- Context Enhanced Transformer for Single Image Object Detection in Video DataSeungjun An, Seonghoon Park, Gyeongnyeon Kim, Jeongyeol Baek et al.AAAI 2024 · 10 citations
- Predict to Detect: Prediction-guided 3D Object Detection using Sequential ImagesSanmin Kim, Youngseok Kim, In-Jae Lee, Dongsuk KumICCV 2023 · 16 citations
- Focus on the Positives: Self-Supervised Learning for Biodiversity MonitoringOmiros Pantazis, Gabriel J. Brostow, Kate E. Jones, Oisin Mac AodhaICCV 2021 · 33 citations
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
