Towards Viewpoint-Invariant Visual Recognition via Adversarial Training
Shouwei Ruan, Yinpeng Dong, Hang Su, Jianteng Peng, Ning Chen, Xingxing Wei
Abstract
Visual recognition models are not invariant to viewpoint changes in the 3D world, as different viewing directions can dramatically affect the predictions given the same object. Although many efforts have been devoted to making neural networks invariant to 2D image translations and rotations, viewpoint invariance is rarely investigated. As most models process images in the perspective view, it is challenging to impose invariance to 3D viewpoint changes based only on 2D inputs. Motivated by the success of adversarial training in promoting model robustness, we propose Viewpoint-Invariant Adversarial Training (VIAT) to improve viewpoint robustness of common image classifiers. By regarding viewpoint transformation as an attack, VIAT is formulated as a minimax optimization problem, where the inner maximization characterizes diverse adversarial viewpoints by learning a Gaussian mixture distribution based on a new attack GMVFool, while the outer minimization trains a viewpointinvariant classifier by minimizing the expected loss over the worst-case adversarial viewpoint distributions. To further improve the generalization performance, a distribution sharing strategy is introduced leveraging the transferability of adversarial viewpoints across objects. Experiments validate the effectiveness of VIAT in improving the viewpoint robustness of various image classifiers based on the diversity of adversarial viewpoints generated by GMVFool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 923a8d40-0702-46ac-8b11-2895347b7832Cited by top-tier papers13
- When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation AttackHanqing Liu, Shouwei Ruan, Yao Huang, Shiji Zhao et al.ICCV 2025 · 13 citations
- Not All Views Are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation ModelsMateusz Michalkiewicz, Sheena Bai, Mahsa Baktashmotlagh, Varun Jampani et al.ICCV 2025 · 8 citations
- Embodied Laser Attack: Leveraging Scene Priors to Achieve Agent-based Robust Non-contact AttacksYitong Sun, Yao Huang, Xingxing WeiACM MM 2024 · 2 citations
- CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language DetectionZhipeng Liu, Chunbo LuoCVPR 2026 · 1 citation
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang et al.ICCV 2025 · 1 citation
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial ViewpointsYinpeng Dong, Shouwei Ruan, Hang Su, Caixin Kang et al.NeurIPS 2022 · 72 citations
- On the Connection between Invariant Learning and Adversarial Training for Out-of-Distribution GeneralizationShiji Xin, Yifei Wang, Jingtong Su, Yisen WangAAAI 2023 · 14 citations
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 61 citations
- Encoding Robustness to Image Style via Adversarial Feature PerturbationsManli Shu, Zuxuan Wu, Micah Goldblum, Tom GoldsteinNeurIPS 2021 · 23 citations
- 360-Attack: Distortion-Aware Perturbations from Perspective-ViewsYunjian Zhang, Yanwei Liu, Jinxia Liu, Jingbo Miao et al.CVPR 2022 · 4 citations
