LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
Er Jin, Qihui Feng, Yongli Mou, Gerhard Lakemeyer, Stefan Decker, Oliver Simons, Johannes Stegmaier
摘要
Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content. This capability is essential in applications such as industrial inspection, where logical anomaly detection (AD) is critical for maintaining high-quality standards and minimizing costly recalls. Previous research in AD has relied on prior knowledge for designing algorithms, which often requires extensive manual annotation effort, significant computing power, and large amounts of data for training. Autoregressive, multimodal Vision Language Models (AVLMs) offer a promising alternative due to their exceptional performance in visual reasoning across various domains. Despite this, their application in logical AD remains unexplored. In this work, we investigate using AVLMs for logical AD and demonstrate that they are well-suited to the task. Combining AVLMs with format embedding and a logic reasoner, we achieve state-of-the-art (SOTA) AD performance on public benchmarks, MVTec LOCO AD, with an AUROC of 86.0% and an F1-max of 83.7% along with explanations of the anomalies. This significantly outperforms the existing SOTA method by 18.1% in AUROC and 4.6% in F1-max score. The dataset, code and supplementary materials are available at https://jasonjin34.github.io/logicad.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionYanhui Li, Yunkang Cao, Chengliang Liu, Yuan Xiong 等AAAI 2026 · 被引用 12 次
- SALAD - Semantics-Aware Logical Anomaly DetectionMatic Fucka, Vitjan Zavrtanik, Danijel SkocajICCV 2025 · 被引用 9 次
- UniAD: Integrating Geometric and Semantic Cues for Unified Anomaly DetectionXiaodong Wang, Hongmin Hu, Fei Yan, Junwen Lu 等ACM MM 2025 · 被引用 2 次
- BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship DetectionMelissa Schween, Mathis Kruse, Bodo RosenhahnCVPR 2026
- LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly DetectionZhichao Zeng, Jiasheng Zhang, Jiyun Sun, Jiangtao Cui 等CVPR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- Logical Anomaly Detection with Text-based Logic via Component-Aware Contrastive Language-Image TrainingSeungeon Lee, Soopil Kim, Sion An, Sang-Chul Lee 等KDD 2025
- Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly DetectionSoopil Kim, Sion An, Philip Chikontwe, Myeongkyun Kang 等AAAI 2024 · 被引用 48 次
- Reasoning-Driven Anomaly Detection and Localization with Image-Level SupervisionYizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang 等CVPR 2026 · 被引用 4 次
- SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentGuoxin Zang, Xue Li, Donglin Di, Lanshun Nie 等ACM MM 2025 · 被引用 8 次
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao 等AAAI 2026 · 被引用 2 次
