DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting
Gaowei Zhang, Lihe Zhang
Abstract
Zero-shot anomaly detection (ZSAD) aims to utilize auxiliary data to train models for generalized learning of unseen categories, which has important application value in fields such as industrial quality inspection and medical diagnosis. Although methods based on CLIP show potential, their pre-training objective of focusing on overall semantic alignment between images and text makes the model insensitive to local details, which is inherently contradictory to the need for fine-grained local features in anomaly detection. Existing improvement methods rely on predefined text prompt frameworks to perceive local information, but struggle to effectively address the issue of insufficient local perception. To address this, this paper proposes a dynamic local visual prompting method based on CLIP (DLVP-CLIP). DLVP dynamically identifies and extracts local visual features from key regions in images as prompt tokens using the Semantic-Aware Local Feature Selector (SLFS) module, and utilizes the multi-modal local prompt (MLoP) module to jointly optimize representations in both visual and textual spaces, achieving more precise cross-modal alignment. Additionally, the high-low frequency decomposition module (HFD) is introduced to separate and process global structural and local textural information via wavelet transformation, thereby enhancing detail perception. Extensive experiments on 13 anomaly detection datasets demonstrate that DLVP-CLIP achieves outstanding ZSAD performance on datasets from the industrial and medical domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 712bdb07-ddc4-4bbe-a065-1d1e5a61ebedBuilds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
- AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly DetectionQihang Zhou, Guansong Pang, Yu Tian, Shibo He et al.ICLR 2024 · 380 citations
- Self-regulating Prompts: Foundational Model Adaptation without ForgettingMuhammad Uzair Khattak, Syed Talal Wasim, Muzammal Naseer, Salman Khan et al.ICCV 2023 · 365 citations
Related papers
- AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP AdaptationQingqing Fang, Wenxi Lv, Qinliang SuACM MM 2025 · 17 citations
- Aligning and Prompting Anything for Zero-Shot Generalized Anomaly DetectionJitao Ma, Weiying Xie, Hangyu Ye, Daixun Li et al.AAAI 2025 · 3 citations
- PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt MixturesYuheng Shao, Lizhang Wang, Changhao Li, Peixian Chen et al.AAAI 2026
- FE-CLIP: Frequency Enhanced CLIP Model for Zero-Shot Anomaly Detection and SegmentationTao Gong, Qi Chu, Bin Liu, Wei Zhou et al.ICCV 2025 · 4 citations
- Bayesian Prompt Flow Learning for Zero-Shot Anomaly DetectionZhen Qu, Xian Tao, Xinyi Gong, Shichen Qu et al.CVPR 2025
