JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
Hyunju Kang, Woohyun Lee, Jaewon Kim, Hogun Park
摘要
Industrial anomaly detection has been significantly advanced by Large Multimodal Models (LMMs), enabling diverse human instructions beyond detection, particularly through visually grounded reasoning for better image understanding. However, LMMs lack domain-specific knowledge, which limits their ability to generate accurate responses in complex industrial scenarios. In this work, we present JUDO, Juxtaposed Domain-Oriented Multimodal Reasoner, a framework that efficiently incorporates domain knowledge and context in visual and textual reasoning. Through visual reasoning, our model segments the defect region by juxtaposing query images with normal images as visual domain context, enabling a fine-grained visual comparative inspection. Furthermore, we inject domain knowledge through supervised fine-tuning (SFT) to enhance context understanding and subsequently guide domain reasoning through reinforcement learning (GRPO) with tailored rewards, opting for a domain-oriented reasoning process. Experimental results demonstrate that JUDO achieves superior performance on the MMAD benchmark, surpassing models such as Qwen2.5-VL-7B and GPT-4o. These results highlight the importance of enhancing domain knowledge and context for effective reasoning in anomaly understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf 等CVPR 2022 · 被引用 1,301 次
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language ModelsZhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 等AAAI 2024 · 被引用 312 次
- Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly DetectionChengjie Wang, Wenbing Zhu, Bin-Bin Gao, Zhenye Gan 等CVPR 2024 · 被引用 71 次
- OntoTune: Ontology-Driven Self-training for Aligning Large Language ModelsZhiqiang Liu, Chengtao Gan, Junjie Wang, Yichi Zhang 等WWW 2025 · 被引用 11 次
相关 Paper
- MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language ModelsXincheng Yao, Zefeng Qian, Chao Shi, Jiayang Song 等CVPR 2026 · 被引用 2 次
- AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward OptimizationJingyi Liao, Yongyi Su, Rong-Cheng Tu, Zhao Jin 等AAAI 2026
- MAU-GPT: Enhancing Multi-type Industrial Anomaly Understanding via Anomaly-aware and Generalist Experts AdaptationZhuonan Wang, Zhenxuan Fan, Siwen Tan, Yu Zhong 等AAAI 2026 · 被引用 1 次
- SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentGuoxin Zang, Xue Li, Donglin Di, Lanshun Nie 等ACM MM 2025 · 被引用 8 次
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly DetectionXi Jiang, Jian Li, Hanqiu Deng, Yong Liu 等ICLR 2025 · 被引用 3 次
