Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
Guangyao Zhai, Yue Zhou, Xinyan Deng, Lars Heckler-Kram, Nassir Navab, Benjamin Busam
Abstract
Few-shot anomaly detection streamlines and simplifies industrial safety inspection. However, limited samples make accurate differentiation between normal and abnormal features challenging, and even more so under category-agnostic conditions. Large-scale pre-training of foundation visual encoders has advanced many fields, as the enormous quantity of data helps to learn the general distribution of normal images. We observe that the anomaly amount in an image directly correlates with the difference in the learnt embeddings and utilize this to design a few-shot anomaly detector termed FoundAD. This is done by learning a nonlinear projection operator onto the natural image manifold. The simple operator acts as an effective tool for anomaly detection to characterize and identify out-of-distribution regions in an image. Extensive experiments show that our approach supports multi-class detection and achieves competitive performance compared to other approaches, while surpassing them in model size and inference efficiency. Backed up by evaluations with multiple foundation encoders, including fresh DINOv3, we believe this idea broadens the perspective on foundation features and advances the field of few-shot anomaly detection. Our code is at https://github.com/ymxlzgy/FoundAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 178765ca-4ae8-4bc3-8106-3e35ddaa6a6fBuilds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace ModelingCamile Lendering, Erkut Akdag, Egor BondarauCVPR 2026 · 12 citations
- FastRecon: Few-shot Industrial Anomaly Detection via Fast Feature ReconstructionZheng Fang, Xiaoyang Wang, Haocheng Li, Jiejie Liu et al.ICCV 2023 · 100 citations
- AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion ModelTeng Hu, Jiangning Zhang, Ran Yi, Yuzhen Du et al.AAAI 2024 · 175 citations
- A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly DetectionShelly Sheynin, Sagie Benaim, Lior WolfICCV 2021 · 106 citations
- FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly DetectionYufei Li, Long Tian, Yuyang Dai, Wenchao Chen et al.CVPR 2026 · 7 citations
