Open-World Object Counting in Videos
Niki Amini-Naieni, Andrew Zisserman
摘要
We introduce a new task of open-world object counting in videos: given a text description, or an image example, that specifies the target object, the objective is to enumerate all the unique instances of the target objects in the video. This task is especially challenging in crowded scenes with occlusions and objects of similar appearance, where avoiding double counting and identifying reappearances is crucial. To this end, we make the following contributions: we introduce a model, CountVid, for this task. It leverages an image-based counting model, and a promptable video segmentation and tracking model, to enable automated open-world object counting across video frames. To evaluate its performance, we introduce VideoCount, a new dataset for this novel task built from the TAO and MOT20 tracking datasets, as well as from videos of penguins and metal alloy crystallization captured by x-rays. Using this dataset, we demonstrate that CountVid provides accurate object counts, and significantly outperforms strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 被引用 14 次
- Referring Multiple Regions with Large Multimodal Models via Contextual Latent SteeringYun Xing, Hanyuan Liu, Jiahao Nie, Shijian LuICML 2026
它引用的顶会 Paper21
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 被引用 482 次
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 被引用 96 次
- A Low-Shot Object Counting Network With Iterative Prototype AdaptationNikola Ðukic, Alan Lukezic, Vitjan Zavrtanik, Matej KristanICCV 2023 · 被引用 91 次
- CLIP-Count: Towards Text-Guided Zero-Shot Object CountingRuixiang Jiang, Lingbo Liu, Changwen ChenACM MM 2023 · 被引用 78 次
相关 Paper
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 被引用 151 次
- VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object TrackingZekun Qian, Ruize Han, Junhui Hou, Linqi Song 等ICCV 2025 · 被引用 3 次
- Video OWL-ViT: Temporally-consistent open-world localization in videoGeorg Heigold, Daniel Keysers, Matthias Minderer, Mario Lucic 等ICCV 2023 · 被引用 22 次
- OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and CaptioningAnwesa Choudhuri, Girish Chowdhary, Alexander G. SchwingNeurIPS 2024 · 被引用 7 次
- UniVS: Unified and Universal Video Segmentation with Prompts as QueriesMinghan Li, Shuai Li, Xindong Zhang, Lei ZhangCVPR 2024
