Do Image Classifiers Generalize Across Time?
Vaishaal Shankar, Achal Dave, Rebecca Roelofs, Deva Ramanan, Benjamin Recht, Ludwig Schmidt
Abstract
Vision models notoriously flicker when applied to videos: they correctly recognize objects in some frames, but fail on perceptually similar, nearby frames. In this work, we systematically analyze the robustness of image classifiers to such temporal perturbations in videos. To do so, we construct two new datasets, ImageNet-Vid-Robust and YTBB-Robust, containing a total of 57,897 images grouped into 3,139 sets of perceptually similar images. Our datasets were derived from ImageNet-Vid and Youtube-BB, respectively, and thoroughly re-annotated by human experts for image similarity. We evaluate a diverse array of classifiers pre-trained on ImageNet and show a median classification accuracy drop of 16 and 10 points, respectively, on our two datasets. Additionally, we evaluate three detection models and show that natural perturbations induce both classification as well as localization errors, leading to a median drop in detection mAP of 14 points. Our analysis demonstrates that perturbations occurring naturally in videos pose a substantial and realistic challenge to deploying convolutional neural networks in environments that require both reliable and low-latency predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 922735da-085f-427c-91e5-a41b2453349bCited by top-tier papers25
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Accuracy on the Line: on the Strong Correlation Between Out-of-Distribution and In-Distribution GeneralizationJohn Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa et al.ICML 2021 · 323 citations
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer et al.NeurIPS 2021 · 304 citations
- Probable Domain Generalization via Quantile Risk MinimizationCian Eastwood, Alexander Robey, Shashank Singh, Julius von Kügelgen et al.NeurIPS 2022 · 99 citations
Related papers
- A Large-Scale Robustness Analysis of Video Action Recognition ModelsMadeline Chantry Schiappa, Naman Biyani, Prudvi Kamtam, Shruti Vyas et al.CVPR 2023
- Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsRunhao Zeng, Xiaoyong Chen, Jiaming Liang, Huisi Wu et al.CVPR 2024 · 6 citations
- ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectChenshuang Zhang, Fei Pan, Junmo Kim, In So Kweon et al.CVPR 2024 · 9 citations
- Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image ModelsHee-Seon Kim, Minji Son, Minbeom Kim, Myung-Joon Kwon et al.ICCV 2023 · 13 citations
- Self-supervised video pretraining yields robust and more human-aligned visual representationsNikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. HénaffNeurIPS 2023 · 27 citations
