AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
Qingqing Fang, Wenxi Lv, Qinliang Su
摘要
Visual anomaly detection has been widely used in industrial inspection and medical diagnosis. Existing methods typically demand substantial training samples, limiting their utility in zero-/few-shot scenarios. While recent efforts have leveraged CLIP's zero-shot recognition capability for this task, they often ignore optimizing visual features to focus on local anomalies, reducing their efficacy. In this work, we propose AF-CLIP (Anomaly-Focused CLIP) by dramatically enhancing its visual representations to focus on local defects. Our approach introduces a lightweight adapter that emphasizes anomaly-relevant patterns in visual features, simultaneously optimizing both class-level features for image classification and patch-level features for precise localization. To capture anomalies of different sizes and improve detection accuracy, prior to the adapter, we develop a multi-scale spatial aggregation mechanism to effectively consolidate neighborhood context. Complementing these visual enhancements, we design learnable textual prompts that generically characterize normal and abnormal states. After optimization on auxiliary datasets using a composite objective function, AF-CLIP demonstrates strong zero-shot detection capability. Our method is also extended to few-shot scenarios by extra memory banks. Experimental results across diverse industrial and medical datasets demonstrate the effectiveness and generalization of our proposed method. Code is available at https://github.com/Faustinaqq/AF-CLIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MRAD: Zero-Shot Anomaly Detection with Memory-Driven RetrievalChaoran Xu, Chengkan Lv, Qiyu Chen, Feng Zhang 等ICLR 2026 · 被引用 9 次
- FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background DisentanglementMing Hu, Yongsheng Huo, Mingyu Dou, Jianfu Yin 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual PromptingGaowei Zhang, Lihe ZhangCVPR 2026
- AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly DetectionQihang Zhou, Guansong Pang, Yu Tian, Shibo He 等ICLR 2024 · 被引用 380 次
- WinCLIP: Zero-/Few-Shot Anomaly Classification and SegmentationJongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang 等CVPR 2023
- FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and SegmentationTao Gong, Qi Chu, Bin Liu, Wei Zhou 等ICCV 2025 · 被引用 4 次
- SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly DetectionChenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu 等ACM MM 2024 · 被引用 10 次
