LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction
Er Jin, Qihui Feng, Yongli Mou, Gerhard Lakemeyer, Stefan Decker, Oliver Simons, Johannes Stegmaier
Abstract
Logical image understanding involves interpreting and reasoning about the relationships and consistency within an image's visual content. This capability is essential in applications such as industrial inspection, where logical anomaly detection (AD) is critical for maintaining high-quality standards and minimizing costly recalls. Previous research in AD has relied on prior knowledge for designing algorithms, which often requires extensive manual annotation effort, significant computing power, and large amounts of data for training. Autoregressive, multimodal Vision Language Models (AVLMs) offer a promising alternative due to their exceptional performance in visual reasoning across various domains. Despite this, their application in logical AD remains unexplored. In this work, we investigate using AVLMs for logical AD and demonstrate that they are well-suited to the task. Combining AVLMs with format embedding and a logic reasoner, we achieve state-of-the-art (SOTA) AD performance on public benchmarks, MVTec LOCO AD, with an AUROC of 86.0% and an F1-max of 83.7% along with explanations of the anomalies. This significantly outperforms the existing SOTA method by 18.1% in AUROC and 4.6% in F1-max score. The dataset, code and supplementary materials are available at https://jasonjin34.github.io/logicad.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f5a6772-9ce7-4869-be32-0ae760bdeeceCited by top-tier papers5
- IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionYanhui Li, Yunkang Cao, Chengliang Liu, Yuan Xiong et al.AAAI 2026 · 12 citations
- SALAD - Semantics-Aware Logical Anomaly DetectionMatic Fucka, Vitjan Zavrtanik, Danijel SkocajICCV 2025 · 9 citations
- UniAD: Integrating Geometric and Semantic Cues for Unified Anomaly DetectionXiaodong Wang, Hongmin Hu, Fei Yan, Junwen Lu et al.ACM MM 2025 · 2 citations
- BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship DetectionMelissa Schween, Mathis Kruse, Bodo RosenhahnCVPR 2026
- LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly DetectionZhichao Zeng, Jiasheng Zhang, Jiyun Sun, Jiangtao Cui et al.CVPR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Logical Anomaly Detection with Text-based Logic via Component-Aware Contrastive Language-Image TrainingSeungeon Lee, Soopil Kim, Sion An, Sang-Chul Lee et al.KDD 2025
- Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly DetectionSoopil Kim, Sion An, Philip Chikontwe, Myeongkyun Kang et al.AAAI 2024 · 48 citations
- Reasoning-Driven Anomaly Detection and Localization with Image-Level SupervisionYizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang et al.CVPR 2026 · 4 citations
- SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentGuoxin Zang, Xue Li, Donglin Di, Lanshun Nie et al.ACM MM 2025 · 8 citations
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao et al.AAAI 2026 · 2 citations
