OV-VOD: Open-Vocabulary Video Object Detection
Zhihong Zheng, Yang Cao, Junlong Gao, Hanzi Wang
摘要
Traditional Video Object Detection (VOD) is limited by pre-defined closed-set categories, restricting its ability to detect novel objects in real-world scenarios. To address this limitation, we make three key contributions. First, we formally define Open-Vocabulary Video Object Detection (Open-Vocabulary VOD) as the task of detecting objects in video streams from open-set categories, including novel categories unseen during training. Second, we establish an evaluation benchmark by utilizing existing datasets (LV-VIS, BURST, and TAO) to bridge the data gap for this new task. Third, we propose OV-VOD, an Open-Vocabulary VOD method that detects objects in videos beyond pre-defined training categories and addresses the shortcomings of image-level open-vocabulary detectors, which generally neglect the essential temporal and spatial information. Specifically, we design a Semantic-Presence Memory Tracking (SPMT) module that propagates object features across frames through a memory bank to leverage temporal consistency. Moreover, we propose a Spatial Object Relationship Distillation loss (L SR ) that captures inter-object spatial dependencies and enhances knowledge transfer during feature distillation. Experiments on multiple video datasets demonstrate that our OV-VOD exhibits superior zero-shot generalization capability compared to existing image-level open-vocabulary object detection methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan 等ICCV 2023 · 被引用 35 次
- Towards Open-Vocabulary Video Instance SegmentationHaochen Wang, Xiaolong Jiang, Xu Tang, Yao Hu 等ICCV 2023 · 被引用 56 次
- OVTrack: Open-Vocabulary Multiple Object TrackingSiyuan Li, Tobias Fischer, Lei Ke, Henghui Ding 等CVPR 2023
- CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object DetectionShiyuan Ma, Donglin Qian, Kai Ye, Shengchuan ZhangAAAI 2025 · 被引用 8 次
- Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object DetectionJiaming Li, Jiacheng Zhang, Jichang Li, Ge Li 等CVPR 2024
