VmAP: A Fair Metric for Video Object Detection
Anupam Sobti, Vaibhav Mavi, M. Balakrishnan, Chetan Arora
Abstract
Video object detection is the task of detecting objects in a sequence of frames, typically, with a significant overlap in content among consecutive frames. Mean Average Precision (mAP) was originally proposed for evaluating object detection techniques in independent frames, but has been used for evaluating video based object detectors as well. This is undesirable since the average precision over all frames masks the biases that a certain object detector might have against certain types of objects depending on the number of frames for which the object is present in a video sequence. In this paper we show several disadvantages of mAP as a metric for evaluating video based object detection. Specifically, we show that: (a) some object detectors could be severely biased against some specific kind of objects, such as small, blurred, or low contrast objects, and such differences may not reflect in mAP based evaluation, (b) operating a video based object detector at the best frame based precision/recall value (high F1 score) may lead to many false positives without a significant increase in the number of objects detected. (c) mAP does not take into account that tracking can be potentially used to recover missed detections in the temporal neighborhood while this can be account for while evaluating detectors. As an alternate, we suggest a novel evaluation metric (VmAP) which takes the focus away from evaluating detections on every frame. Unlike mAP, VmAP rewards a high recall of different object views throughout the video. We form sets of bounding boxes having similar views of an object in a temporal neighborhood and use a set-level recall for evaluation. We show that VmAP is able to address all the challenges with the mAP listed above. Our experiments demonstrate hidden biases in object detectors, shows upto 99% reduction in false positives while maintaining similar object recall and shows a 9% improvement in correlation with post-tracking performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e08fb42-3939-47a2-90ff-0031f4672143Builds on6
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- A Delay Metric for Video Object Detection: What Average Precision Fails to TellHuizi Mao, Xiaodong Yang, Bill DallyICCV 2019 · 44 citations
- CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object DetectionZhiwei Dong, Guoxuan Li, Yue Liao, Fei Wang et al.CVPR 2020
Related papers
- Optimal Correction Cost for Object Detection EvaluationMayu Otani, Riku Togashi, Yuta Nakashima, Esa Rahtu et al.CVPR 2022 · 18 citations
- Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video RetrievalYang Liu, Qianqian Xu, Peisong Wen, Siran Dai et al.ACM MM 2024 · 9 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Margin MCC: Chance-Robust Metric for Video Boundary Detection with Allowed MarginKosuke Mizufune, Shunsuke Tanaka, Toshihide Yukitake, Tatsushi MatsubayashiACM MM 2023
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
