Progress and Limitations of Deep Networks to Recognize Objects in Unusual Poses
Amro Abbas, Stéphane Deny
Abstract
Deep networks should be robust to rare events if they are to be successfully deployed in high-stakes real-world applications (e.g., self-driving cars). Here we study the capability of deep networks to recognize objects in unusual poses. We create a synthetic dataset of images of objects in unusual orientations, and evaluate the robustness of a collection of 38 recent and competitive deep networks for image classification. We show that classifying these images is still a challenge for all networks tested, with an average accuracy drop of 29.5% compared to when the objects are presented upright. This brittleness is largely unaffected by various network design choices, such as training losses (e.g., supervised vs. self-supervised), architectures (e.g., convolutional networks vs. transformers), dataset modalities (e.g., images vs. image-text pairs), and data-augmentation schemes. However, networks trained on very large datasets substantially outperform others, with the best network tested-Noisy Student EfficentNet-L2 trained on JFT-300M-showing a relatively small accuracy drop of only 14.5% on unusual poses. Nevertheless, a visual inspection of the failures of Noisy Student reveals a remaining gap in robustness with the human visual system. Furthermore, combining multiple object transformations-3D-rotations and scaling-further degrades the performance of all networks. Altogether, our results provide another measurement of the robustness of deep networks that is important to consider when using them in the real world. Code and datasets are available at https://github.com/amro- kamal/ObjectPose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Does Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark IbrahimICLR 2024 · 2 citations
- A Deep Learning Model of Mental Rotation Informed by Interactive VR ExperimentsRaymond Khazoum, Daniela Fernandes, Aleksandr Krylov, Qin Li et al.ICML 2026 · 2 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Contemplating Real-World Object ClassificationAli BorjiICLR 2021 · 2 citations
- Transformation-Equivariant 3D Object Detection for Autonomous DrivingHai Wu, Chenglu Wen, Wei Li, Xin Li et al.AAAI 2023 · 158 citations
- Rotationally Equivariant 3D Object DetectionHong-Xing Yu, Jiajun Wu, Li YiCVPR 2022 · 31 citations
- ConDor: Self-Supervised Canonicalization of 3D Pose for Partial ShapesRahul Sajnani, Adrien Poulenard, Jivitesh Jain, Radhika Dua et al.CVPR 2022 · 28 citations
- ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial ViewpointsYinpeng Dong, Shouwei Ruan, Hang Su, Caixin Kang et al.NeurIPS 2022 · 72 citations
