Progress and Limitations of Deep Networks to Recognize Objects in Unusual Poses
Amro Abbas, Stéphane Deny
摘要
Deep networks should be robust to rare events if they are to be successfully deployed in high-stakes real-world applications (e.g., self-driving cars). Here we study the capability of deep networks to recognize objects in unusual poses. We create a synthetic dataset of images of objects in unusual orientations, and evaluate the robustness of a collection of 38 recent and competitive deep networks for image classification. We show that classifying these images is still a challenge for all networks tested, with an average accuracy drop of 29.5% compared to when the objects are presented upright. This brittleness is largely unaffected by various network design choices, such as training losses (e.g., supervised vs. self-supervised), architectures (e.g., convolutional networks vs. transformers), dataset modalities (e.g., images vs. image-text pairs), and data-augmentation schemes. However, networks trained on very large datasets substantially outperform others, with the best network tested-Noisy Student EfficentNet-L2 trained on JFT-300M-showing a relatively small accuracy drop of only 14.5% on unusual poses. Nevertheless, a visual inspection of the failures of Noisy Student reveals a remaining gap in robustness with the human visual system. Furthermore, combining multiple object transformations-3D-rotations and scaling-further degrades the performance of all networks. Altogether, our results provide another measurement of the robustness of deep networks that is important to consider when using them in the real world. Code and datasets are available at https://github.com/amro- kamal/ObjectPose.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Does Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark IbrahimICLR 2024 · 被引用 2 次
- A Deep Learning Model of Mental Rotation Informed by Interactive VR ExperimentsRaymond Khazoum, Daniela Fernandes, Aleksandr Krylov, Qin Li 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Contemplating Real-World Object ClassificationAli BorjiICLR 2021 · 被引用 2 次
- Transformation-Equivariant 3D Object Detection for Autonomous DrivingHai Wu, Chenglu Wen, Wei Li, Xin Li 等AAAI 2023 · 被引用 158 次
- Rotationally Equivariant 3D Object DetectionHong-Xing Yu, Jiajun Wu, Li YiCVPR 2022 · 被引用 31 次
- ConDor: Self-Supervised Canonicalization of 3D Pose for Partial ShapesRahul Sajnani, Adrien Poulenard, Jivitesh Jain, Radhika Dua 等CVPR 2022 · 被引用 28 次
- ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial ViewpointsYinpeng Dong, Shouwei Ruan, Hang Su, Caixin Kang 等NeurIPS 2022 · 被引用 72 次
