DE-CLIP: Few-Shot Anomaly Detection via Difference-Guided Embedding Editing
Yage Zhang, Yukun Jiang, Michael Backes, Yang Zhang
摘要
Anomaly detection (AD) plays a critical role in applications such as automated industrial inspection and medical image analysis. Empowered by the strong pre-trained vision-language model, CLIP, recent years have witnessed the emergence of several CLIP-based few-shot AD methods. Due to the overlap between the embedding distributions of normal and anomalous samples, many existing approaches introduce additional model training for more discriminative text embeddings. However, we demonstrate that such training is not necessary. Specifically, we find that this embedding overlap can be separated by introducing a Differenceguided vector for embedding Editing (DiffEdit). Based on this finding, we propose DE-CLIP, a simple yet effective framework based on DiffEdit, which directly edits text embeddings based on the textual and visual differences between normal and anomalous samples, resulting in more discriminative embeddings for AD. Extensive experiments on industrial and medical datasets demonstrate the superiority of our proposed DE-CLIP compared with existing baselines. For instance, on the MVTec dataset, DE-CLIP achieves 96.6% and 96.7% AUROC on anomaly classification and segmentation, surpassing both training-based and training-free methods. In addition, we observe that introducing DiffEdit into other trainingfree baselines could also significantly improve their performance, highlighting the potential of DiffEdit to promote better AD. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- WinCLIP: Zero-/Few-Shot Anomaly Classification and SegmentationJongheon Jeong, Yang Zou, Taewan Kim, Dongqing Zhang 等CVPR 2023
- AdaptCLIP: Adapting CLIP for Universal Visual Anomaly DetectionBin-Bin Gao, Yue Zhou, Jiangtao Yan, Yuezhi Cai 等AAAI 2026 · 被引用 21 次
- SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly DetectionChenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu 等ACM MM 2024 · 被引用 10 次
- AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly DetectionQihang Zhou, Guansong Pang, Yu Tian, Shibo He 等ICLR 2024 · 被引用 380 次
- AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP AdaptationQingqing Fang, Wenxi Lv, Qinliang SuACM MM 2025 · 被引用 17 次
