Test-Time Canonicalization by Foundation Models for Robust Perception
Utkarsh Singhal, Ryan Feng, Stella X. Yu, Atul Prakash
Abstract
Perception in the real world requires robustness to diverse viewing conditions. Existing approaches often rely on specialized architectures or training with predefined data augmentations, limiting adaptability. Taking inspiration from mental rotation in human vision, we propose FOCAL, a testtime robustness framework that transforms the input into the most typical view. At inference-time, FOCAL explores a set of transformed images and chooses the one with the highest likelihood under foundation model priors. This test-time optimization boosts robustness while requiring no retraining or architectural changes. Applied to models like CLIP and SAM, it significantly boosts robustness across a wide range of transformations, including 2D and 3D rotations, contrast and lighting shifts, and day-night changes. We also explore potential applications in active vision. By reframing invariance as a test-time optimization problem, FOCAL offers a general and scalable approach to robustness. Our code is available at: https://github.com/sutkarsh/focal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric NetworksYa-Wei Eileen Lin, Ron LevieICLR 2026 · 4 citations
- Inverting Data Transformations via Diffusion SamplingJinwoo Kim, Sékou-Oumar Kaba, Jiyun Park, Seunghoon Hong et al.ICML 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- MEMO: Test Time Robustness via Adaptation and AugmentationMarvin Zhang, Sergey Levine, Chelsea FinnNeurIPS 2022 · 595 citations
- A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language ModelsMujtaba Hussain Mirza, Antonio D’Orazio, Odelia Melamed, Iacopo MasiCVPR 2026 · 2 citations
- SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language ModelsSunoh Kim, Daeho UmICML 2026
- On the Test-Time Zero-Shot Generalization of Vision-Language Models: Do we Really need Prompt Learning?Maxime Zanella, Ismail Ben AyedCVPR 2024 · 21 citations
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language ModelsZhiwei Li, Yitian Pang, Weining Wang, Zhenan Sun et al.CVPR 2026 · 2 citations
