ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
Kai Zhang, Zekai Zhang, Xihe Sun, Anpeng Wang, Jingmeng Nie, Qinghui Chen, Han Hao, Jianyuan Guo, Jinglin Zhang
摘要
Automatic vision inspection holds significant importance in industry inspection. While multimodal large language models (MLLMs) exhibit strong language understanding capabilities and hold promise for this task, their performance remains significantly inferior to that of human experts. In this context, we identify two key challenges: (i) insufficient integration of anomaly detection (AD) knowledge during pre-training, and (ii) the lack of technically precise and context-aware language generation for anomaly reasoning. To address these issues, we propose ADSeeker, an anomaly task assistant designed to enhance inspection performance through knowledge-grounded reasoning. AD-Seeker first leverages a curated visual document knowledge base, SEEK-MVTec&VisA (SEEK-M&V), which we construct to address the limitations of existing resources that rely solely on unstructured text. SEEK-M&V includes semantic-rich descriptions and image-document pairs, enabling more comprehensive anomaly understanding. To effectively retrieve and utilize this knowledge, we introduce the Query Image-Knowledge Retrieval-Augmented Generation (Q2K RAG) framework. To further enhance the performance in zero-shot anomaly detection (ZSAD), ADSeeker leverages the Hierarchical Sparse Prompt mechanism and type-level features to efficiently extract anomaly patterns. Furthermore, to tackle the challenge of limited industry anomaly detection (IAD) data, we introduce the largestscale AD dataset, Multi-type Anomaly (MulA), encompassing 72 multi-scale defect types across 26 categories. Extensive experiments show that our plug-and-play framework, ADSeeker, achieves state-of-the-art zero-shot performance on several benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly DetectionQihang Zhou, Guansong Pang, Yu Tian, Shibo He 等ICLR 2024 · 被引用 380 次
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 等AAAI 2024 · 被引用 312 次
- FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality LocalizationZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 等ACM MM 2024 · 被引用 52 次
相关 Paper
- Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language ModelsJiacong Xu, Shao-Yuan Lo, Bardia Safaei, Vishal M. Patel 等CVPR 2025
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu 等ICLR 2025 · 被引用 3 次
- MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language ModelsXincheng Yao, Zefeng Qian, Chao Shi, Jiayang Song 等CVPR 2026 · 被引用 2 次
- Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and BaselineZekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding 等AAAI 2025 · 被引用 4 次
- Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly DetectionDahu Shi, Chengshen He, Shaochen Zhang, Bo Qian 等CVPR 2026
